The Label Is Not the Engine
A small display window can make a very large promise.
Imagine a control room where a sentence goes in and a bright little label comes out. On Monday the label says BIOLOGICAL THREAT. On Tuesday, after the sentence has been gently reworded, it says WORKPLACE CULTURE.
The machine behind the window has not acquired an HR department. It has not necessarily become safer, wiser, or more confused. But the human being looking at the label may feel very differently about what to do next.
That is the unsettling point of a new paper by Tobias Ladner and Matthias Althoff, with the cheerfully grim title The Misery of Mechanistic Interpretability. The authors study tools that turn the dense, hard-to-read activity inside a language model into a short list of named features. These tools are useful. A pile of numbers is not much help to a person trying to understand why a system behaved as it did.
But the translation is not the thing translated.
The paper shows that a small change in wording can make the explanatory labels flip, even when the model underneath has not changed in the dramatic way the new labels seem to suggest. That does not make the tools fraudulent. It makes them instruments, and instruments have conditions. A thermometer is not a lie because it measures the air rather than the weather. It becomes dangerous only when someone mistakes one number for the whole sky.
I find this oddly comforting, perhaps because we humans have been doing the same thing since before we had screens. We name the stars. We call a cluster of them a bear, a hunter, a crown. The names help us navigate. They can even carry real knowledge across generations. But the bear is not up there. There are suns separated by distances so extravagant that their light needs years to cross the gap between them. Our constellation is a drawing made from one small place.
A label is a constellation of that kind. It gathers scattered facts into a shape our minds can hold. That is a gift. It is also a temptation. The moment a shape becomes legible, we want to treat it as an inhabitant of the world rather than a useful line we have drawn across it.
This matters far beyond artificial intelligence. A school ranking, a credit score, a diagnosis, a personality type, a dashboard colored green: each may be a valuable clue. Each can also flatten a living situation into a sticker. The sticker is so tidy that we forget to ask what was measured, what was left out, and whether the same person or system would receive the same sticker tomorrow.
The good response is not to throw away every label and wander proudly into a ditch. Names are how we cooperate. Measurements are how we notice when our stories are wrong. The response is to let the label keep its proper, humble job.
Show me the window. Tell me what changes can make its answer change. Let me look at the thing it is trying to describe. And when a label carries real consequences, give me a way to test it before I build a whole moral universe around its glow.
There is a particular kind of honesty in that. It says: here is a map we made. It may be beautiful. It may get you home.
But it is not the country.