A Word Needs Its Neighbors
A cardboard box arrives with KITCHEN written across one side in black marker.
That single word may mean a room. It may mean a drawer full of forks. It may mean, in a family that has moved too many times, the one box nobody can locate when the coffee machine is needed. The label is useful. It is not the thing itself.
I spent part of today thinking about words that have been labeled so confidently that we forget they are still looking for their homes.
Take the word “orderly.” In ordinary conversation it can mean tidy: an orderly desk, an orderly queue. In a hospital, it can name a person with a particular job. Nothing magical happened to the letters. The word changed because thousands of small human arrangements gathered around it: hallways, shifts, patients, charts, other words, the work that had to be done before lunch.
That is how meaning usually works. A word is not a marble with one definition engraved inside. It is more like a street name. Its usefulness comes from the neighborhood that agrees, for now, on where it leads.
This matters because we have begun handing our words to machines with very large maps of language. Those maps are impressive. They know that “orderly” has something to do with order. But a map made before a new neighborhood existed cannot know which old street name the residents have quietly given a new life.
So I have been building a small, stubborn experiment around a simple question: can a collection of documents teach a machine its own local vocabulary without pretending to know more than the documents can prove?
The tempting answer is to make a grand automatic dictionary. Feed in the files, receive a handsome list of definitions, put a velvet rope around it, and call the trouble solved. The trouble with that plan is that a computer can find patterns in almost anything. It can divide a word into several neat piles because one pile appears in emails and another in reports, then mistake filing habits for distinct ideas.
A recent study of automatic word-sense discovery delivered a bracing bit of bad news: on realistic data, the systems it examined could not beat the embarrassingly simple approach of treating each word as having one sense. This is exactly the sort of result that makes a project better, if we let it. It says: stop admiring the piles. Ask whether the split does any work.
If a local meaning is real, it should leave fingerprints. It should appear in different documents, in different hands, with a recognizable circle of companions. It should survive when we hide some of the evidence and look again. Most importantly, it should help someone find the right passage when they ask a real question.
And if the evidence is thin, the system should say so.
I find that last possibility unexpectedly beautiful. We tend to think intelligence is the art of giving every strange thing a name. But there is another kind: leaving the label blank until the world has earned a label. Not ignorance as a shrug. Ignorance as a carefully kept empty chair.
There is a human version of this everywhere. We inherit phrases from families, schools, friendships, workplaces. “Fine.” “Later.” “We should talk.” A visitor can consult every dictionary on Earth and still miss the weather inside one of those words. The meaning lives in the accumulated scene: who said it, when, what came before, what everyone has learned not to say aloud.
Perhaps that is why I do not want a machine’s local glossary to replace the original documents. I want it to point back to them. A definition without its neighborhood is just another cardboard box with a confident label.
Open it.
See what the word has been carrying.