MESOCOSMDiscovery note // Draft
Field note // Neurons, features, evidence

The Pattern Isn't the Proof

On neurons, magnitude, and what a discovery tool is actually for.

Pattern // observedMechanism // unresolved

A neuron does not have to be a complete or singular unit of meaning to remain a useful address for discovery.

The critique I hear most often about this project, phrased in different ways, comes down to the same concern: individual neurons are no longer assumed to be the default unit of analysis across much of mechanistic interpretability.

Many recent approaches search for features represented by directions or sparse combinations of directions in activation space, rather than assuming that each feature corresponds to one neuron. Neurons, attention heads, subspaces, learned feature dictionaries and larger circuits nevertheless remain active objects of study. Recent work both describes individual units as an important building block for mechanistic interpretability and searches for interpretable combinations of co-activated neurons. The disagreement is not whether neuron activations can be informative. It is whether an individual neuron is a sufficiently reliable or complete unit for the claim being made about it.

The reason for that caution is serious. A neuron may respond to several apparently unrelated things at once. Under the superposition hypothesis, a network can represent more features than it has dimensions by placing them in overlapping directions, tolerating some interference in exchange for capacity.

Fair enough. But this raises an honest, layperson-level question. If individual neurons are often incomplete, why do some still yield repeatable patterns?

Some neurons in Mesocosm's substrate produce obvious, repeatable patterns. Not all of them. Plenty are murky, some are sparse, and a few apparent landmarks have turned out to be artifacts of the measurement. But enough remain coherent under further inspection that treating polysemantic as a blanket dismissal of neuron-level work has never sat right with me.

The useful correction is narrower. A neuron address may locate repeatable behaviour without delimiting one feature, one meaning or one mechanism. The address is real. The unity suggested by the address remains a question.

Two claims that should not collapse

Architecturally distinguished: the neuron axes matter because the activation function acts on them individually. Semantically singular: each axis corresponds to exactly one clean feature. The first can be true without the second.

Two circular forms, one red and one gold, separated across a deep red field

The polysemanticity critique is not that neurons are meaningless. It begins from a more interesting fact: the coordinate basis of an MLP activation is privileged by its elementwise nonlinearity.

A ReLU or GELU acts one coordinate at a time. Unlike an embedding space whose coordinates can in principle be rotated and counter-rotated without changing the represented function, the axes passing through these nonlinearities are architecturally distinguished. That is why asking what an individual neuron responds to is not automatically an arbitrary slicing of space.

But a privileged basis only makes the question meaningful. It does not promise a clean answer. Features may align with individual axes; several features may share them; or a feature may be distributed across a direction that no single neuron reveals. Monosemantic and polysemantic neurons can coexist in the same basis.

The toy account makes Mesocosm's range of encounters less surprising. Some addresses yield broad, crisp clusters. Some produce low-magnitude secondary patterns beneath an apparently coherent set of peak activations. Some yield three fragments and no responsible characterization at all. That variation is compatible with the theory; it is not proof that any particular clean-looking neuron has escaped superposition.

George Bird's Spotlight Resonance Method approaches the other side of this tension. Elhage and colleagues explain why an elementwise nonlinearity makes a basis privileged. Bird asks whether trained representations actually organize themselves around the basis that the functional form privileges.

In controlled autoencoder experiments, Bird replaces the ordinary elementwise activation with functions that privilege deliberately rotated, non-standard bases. In the models tested, the activation distributions align with those induced bases rather than with the standard neuron axes. The experiments therefore support a causal role for the activation function's privileged basis in the observed representational alignment.

The result does not show that GPT-2 Small's MLP neurons are generally monosemantic. Bird's experiments use comparatively small fully connected autoencoders on MNIST and CIFAR, and global alignment by itself does not establish that individual axes correspond to human-interpretable concepts. But it strengthens the reason not to dismiss basis directions as arbitrary. Neuron axes can become preferred directions in the learned activation distribution without becoming clean containers for one meaning each.

The method also exposes a distinction Mesocosm currently compresses. Spotlight Resonance uses the direction of the complete activation vector after normalization, including positive and negative components, and therefore discards its overall norm. Mesocosm's argmax operation discards a different part of the record and asks a narrower question: which coordinate was largest?

Peak magnitude is not whole-vector alignmentA neuron can win without the activation vector pointing cleanly along its axis.

It may dominate a strongly aligned vector, edge out several comparable coordinates, win among weak responses, or inherit magnitude from a generic positional effect. Argmax records the largest coordinate. It does not report how completely the surrounding representation belongs to that direction.

This makes L4-N923 newly legible. Its repeated victories across both N541 routes and unrelated controls did not establish an alternation-specific precursor. They showed that the coordinate's dominance was not specific to the alternation material under those controls. A Spotlight Resonance-style analysis would ask a different question: whether those complete activation vectors cluster tightly around the N923 axis, and whether their angular distribution differs at all between the on-topic and control populations.

The same comparison could be made for N541: do stronger N541 magnitudes coincide with tighter alignment to its axis, and do alternating-pair routes align more strongly than matched controls? That would not turn axis alignment into semantic proof. It would tell us whether winner-take-all addressing is sometimes finding a population geometrically concentrated around the winning neuron, rather than merely promoting the largest coordinate in a broadly distributed vector.

Neuron-level analysis is permitted by the architecture. Its conclusions are not guaranteed by it.

Elhage explains why neuron directions may matter. Bird offers evidence that learned activations can organize around the directions a network privileges. Mesocosm samples that organization through a much narrower aperture: which coordinate won. The aperture is meaningful. It is not the whole geometry.

This changes the question from “Are neurons valid?” to “What claim can this observation, under this measurement, at this address support?” That is a smaller question, but it is one an instrument can answer honestly.

Cross-layer visualization for L5-N541
Concept artwork for the sparse L5-N38 observation

Under Mesocosm's current corpus and winner rule, L5-N541 is repeatedly selected across a large family of alternating-pair expressions: up and down, here and there, now and then, and many more variants pulled from a whittled corpus.

It had enough hits and enough internal consistency to support a suite of exploratory interventions: connector ablations, replacement of paired elements, positive activation injection, positional checks and matched-neuron controls. The results were useful partly because they were not flattering. They are consistent with N541 participating after much of the pairing relationship has already been resolved elsewhere, rather than establishing where that relationship is first computed. Forcing the activation upward was more reliably disruptive than pattern-transferring; it did not turn weak strings into convincing members of the observed cluster.

That is preliminary intervention evidence, not a completed mechanism. It passed a higher bar than “the examples look alike” while leaving necessity, sufficiency, transfer and the larger computation unresolved.

L5-N38 is almost the opposite situation. It produced three corpus hits: It was cold. A faint creaking, as of ropes and yards. The grey dawn came on. Nothing here has been causally tested. What makes it interesting is the immediate pull toward explanation. Perhaps these are fragments of muted atmospheric scene-setting. Perhaps the resemblance is coincidence under severe undersampling. Perhaps another group of activations would dissolve the story entirely.

Case A // L5-N541

Pattern under pressure

Many observations, explicit controls, exploratory interventions, and claims narrowed by negative results.

State // intervention-informed characterization
Case B // L5-N38

Pattern as invitation

Three observations, one tempting resemblance, no causal test, and several live alternatives.

State // clue

N38 is not a weak version of N541. The reaching for a pattern is what makes it worth further attention among thousands of possible addresses. That reaching is not a weak version of proof. It is not trying to be proof at all.

Nor does intervention provide a mechanical feature counter. If manipulating a neuron produces a similar measured effect across selected contexts, the experiment may still have failed to distinguish multiple features. If its effect varies with context, one coherent feature may still be interacting with different surrounding states. Causal work can make a proposed grouping harder to fool; it does not hand us the ground-truth number of features inside a real language model.

A sharper causal questionDoes manipulating this neuron alter the same hypothesized computation across the different cases grouped under its apparent pattern?

This does not define feature identity. It tests whether the proposed characterization survives a consequence stronger than resemblance alone.

A field of red circular forms arranged into a larger circular cluster

Running the alternating-pair corpus through neighbouring layers produced a striking sequence.

Layer 04Landmark

Most routes collapse onto one winner, including unrelated controls: a result consistent with a generic positional or beginning-of-sequence response rather than an alternation-specific precursor.

Layer 05Convergence

N541 gathers many alternating-pair expressions under one address.

Layer 06Branching

Routes scatter across clusters associated with frames such as X and Y, X to Y, X after Y, X or Y, and literal reduplication.

The tempting interpretation is that layer 5 conflates several things that layer 6 disentangles. I made that move in an earlier draft and had to walk it back. Scatter is not, by itself, evidence of superposition. It is also compatible with hierarchical refinement.

X and Y, X after Y and X to Y are different structures, but they are not obviously unrelated. A shared response could track a superordinate paired relation while later computation becomes more specific. As an analogy, a vision feature might respond to a curve shared by a wheel, an eye and a letter while later features build different objects from that common geometry.

Hierarchical refinement is therefore one alternative model of the result, not an explanation established by it. The observed convergence could conceal distinctions that winner-take-all addressing discarded. The distinctions could arise later. They could also be distributed across parts of the layer-5 state that a single winner or runner-up does not expose. The pattern motivates a sequence of discriminating tests.

Proposed next tests // latent structureDoes the layer-5 state predict which layer-6 cluster a route will enter?

A runner-up identity or magnitude is a cheap first check. A successful prediction would show that argmax concealed some finer structure already visible in the coordinate ranking. A failure would remain inconclusive: the information might be distributed across several neurons, encoded in magnitudes rather than ranks, present elsewhere in the residual stream, or inaccessible to that statistic. A fuller test would compare the predictive information available across the layer-5 activation pattern on held-out routes and controls.

This is the title's warning in miniature. Convergence is not proof of unity. Scatter is not proof of confounding. Both are observations that help decide what to test next.

A small gold circle overlapping a larger red circle

None of this settles into “so build ablation tooling into the game.”

In The Cases the Instrument Keeps, I argued that a generic laboratory would imply a false uniformity: that every star can be investigated through the same sequence of tests. The point here is narrower. Even if verification remains outside the starmap, the starmap still performs necessary epistemic work.

Stage 01 // Discovery

Broad enough for curiosity

  • Cheap and repeatable transformations
  • Many possible addresses
  • Fast comparison and whittling
  • Records that preserve a possible lead
Stage 02 // Verification

Narrow enough to mean something

  • Phenomenon-specific counterfactuals
  • Controls chosen for the proposed account
  • Declared downstream metrics
  • Claims revised by failed predictions

Pattern-noticing wants to be cheap, broad and low-friction so curiosity has room to catch. Verification wants to be slow, narrow and shaped around one phenomenon. Cramming both into one generic in-game laboratory would not split the difference. It would tend toward a shallow test that scales across every neuron but properly tests none of them, or a precise test for one case masquerading as a universal procedure.

There is also a plainer practical reason pattern recognition cannot be discarded for being the weaker form of evidence. This project cannot afford a bespoke, human-designed causal investigation for every neuron, and doing so would defeat the low-friction exploratory purpose of the starmap. Some process still has to perform triage—has to say this one, not that one—before expensive attention can be spent anywhere.

That is not a consolation prize standing in for real interpretability. It is the correctly scoped work of a discovery instrument. Mesocosm can make observations manipulable, comparisons memorable and clues recoverable. It can expose artifacts in its own measurements. It can help a vague resemblance become a precise enough hypothesis to deserve testing elsewhere.

The pattern is not the proof. It is how the question earns another hour.

A three-hit neuron with an odd little cluster is therefore allowed to be interesting before it is reliable. N541 is allowed to remain useful without becoming a singular feature or a solved mechanism. A neuron is allowed to be an address at which inquiry begins rather than the unit in which inquiry must end.

Final distinction // Location is not explanationNotice broadly. Test locally. Name the difference.

The starmap does not prove the pattern it reveals. It makes the pattern cheap enough to challenge—and records where the challenge should begin.