MESOCOSM //
Field note // literature reflection
31 Aug 2026
Selection // legibility // interpretability

The cases the instrument keeps.

We look at L5-N541 because it yields insight. What does that choice hide about the rest of the starmap?

PART 01 // A PRODUCTIVE RESIDENT AND AN EXPEDITION CHOSEN BY SOMEONE ELSE
A large gold sphere held inside a red field of smaller points, with much of the surrounding field receding into darkness

There are 3,072 neuron addresses in each MLP layer of GPT-2 Small. One of them has become unusually articulate.

L5-N541 repeatedly rewards attention. In the corpus routes collected for it, connected parallel members keep arriving at the point where the second half completes the form: up and down, true or false, one by one, wave after wave. The first observations survived deletion tests. They suggested linguistic controls. Those controls produced failures and refinements. The refinements became interventions, matched comparisons, held-out predictions and a branching account of several downstream destinations.

This history is already documented elsewhere on the site. It does not need to be performed again here. What matters for the present argument is simpler: 541 was worth following because each stage made another stage possible.

There were enough natural examples to see something recur. The recurrence was legible enough to name provisionally. The provisional description was specific enough to challenge. Some challenges failed; others returned differences that could be measured. The object did not become simple, but it remained experimentally generous.

We look at 541 because it yields insight. A true sentence, and therefore a methodological problem

There is nothing inherently dishonest about choosing a productive case. A case study needs a case capable of sustaining study. An instrument is often demonstrated where its operation can be seen clearly. If the question is whether this workflow can sometimes move from an activation winner toward a bounded behavioural characterization and a causal test, 541 is strong evidence that it can.

But a productive example can begin to answer a different question merely by occupying the centre of the picture. How often does an arbitrary neuron permit this kind of progress? How many addresses produce a coherent-enough pattern to justify controls? How many produce six observations, several incompatible impressions, a positional artifact, a token-level coincidence, or nothing under the current measurement protocol at all?

The profile does not claim to answer those questions. The danger is quieter than an explicit overclaim. A detailed resident accumulates images, experiments, captions, links and narrative weight. It becomes the neuron through which the instrument is understood. The surrounding population remains present in principle but disappears from the reader's experience.

Curated depth // L5-N541

A case that keeps opening.

Repeated forms support a provisional description; that description supports controls; the controls support more precise interventions and held-out tests.

Selection
Productive case
Corpus routes
80
Follow-up
V7–V12
Current status
Bounded characterization
External request // N2222 → L5-N2223

A clue that may stop here.

A requested address is absent under the protocol. Its substituted neighbour returns six hits, a narrow resemblance and one resistant outlier.

Selection
Reddit request
Corpus hits
6
Substitution
Explicitly arbitrary
Current status
Clue, not interpretation

A small experiment on Reddit offered a different way to choose where to look. Instead of beginning with a neuron already known to produce an interesting pattern, I asked strangers to supply numbers from 0 to 3071. I would visit the requested neuron—or, when it did not appear under the current measurement protocol, one of its nominal neighbours—and report what I found.

Someone chose 2222. N2222 was not present. The drones returned from L5-N2223 instead.

There were only six corpus hits. After whittling, several routes repeatedly collapsed toward to whelm, was whelmed, had whelmed, by whalemen and the whalemen. A small orthographic or token-level neighbourhood seemed possible. One long outlier would not cooperate. The most responsible description available was not a characterization, but a clue.

The expedition became I Am Underwhalemanned, in part because the Reddit reply produced an impeccable joke. It also preserved something the polished 541 record cannot provide by itself: an encounter selected before anyone knew whether it would be productive.

This does not make N2223 representative. One crowd-supplied number is not a survey. The requested address was missing, the substitution rule was arbitrary, and six hits cannot reveal the distribution of interpretability across a layer. But the episode changes the texture of the archive. It records the possibility that an honest expedition returns with very little: a resemblance, an outlier and no justified next experiment.

N2223 is not a disappointing version of N541. It answers a different question. The 541 record asks how far sustained attention can sometimes travel. The Reddit expedition asks what happens when the destination is chosen before its promise is known.

The first distinction

A case can be exemplary without being representative. The trouble begins when the brightness of the example erases the field from which it was selected.

Abstract red and gold composition

Cherry-picking sounds like something that happens at the end: the investigator gathers the results, hides the awkward ones and publishes a highlight reel.

This reflection grew out of reading Leonard Bereska and Efstratios Gavves' broad survey, Mechanistic Interpretability for AI Safety: A Review. Several of its challenge categories gave names to problems Mesocosm had already encountered in practice; others exposed pressures the project had not yet made explicit. What follows is not a summary of the survey. It is an attempt to let its objections act on this particular instrument.

That version exists. It is also too simple for the problem here. Long before a neuron receives a profile, Mesocosm has already decided what can appear as an observation. The corpus supplies some strings and not others. The measurement protocol compresses a field of activations into a winner. Only neurons that win somewhere enter the available population. Whittling keeps deletions that preserve a declared response. The investigator notices some regularities before others. A provisional description makes certain controls imaginable. An experiment is designed for the phenomena that remain legible long enough to ask a causal question.

Publication is only the final aperture.

In their section on “Cherry-Picking and Streetlight Interpretability,” Bereska and Gavves describe a field drawn toward conditions of maximal interpretability: small models, toy tasks and convincing examples that can be made unusually clear. Their warning arrives in a review whose stated ambition is a granular, causal understanding of learned mechanisms. The distance between that ambition and the cases that can currently be investigated is part of the problem. We learn most where our instruments can already see.

Mesocosm inherits that problem at two scales. GPT-2 Small is itself a tractable model: old, comparatively small, downloadable, reproducible on ordinary hardware and unusually available for this kind of personal investigation. Then, inside GPT-2 Small, the project is drawn toward neurons such as 541, whose behaviour can be made coherent enough to continue.

Neither choice invalidates the work. Together they define its illuminated area.

Selection 01

Corpus

Only phenomena expressed in the available strings can become natural observations.

Selection 02

Winner

Argmax promotes one neuron and discards most of the surrounding activation structure.

Selection 03

Survival

Whittling retains textual reductions that preserve the response already chosen for study.

Selection 04

Attention

Some clusters become perceptible, memorable and linguistically describable sooner than others.

Selection 05

Experiment

Some phenomena admit useful counterfactuals, controls and outcome measures; others resist them.

Selection 06

Archive

A small subset becomes profiles, writeups, images and the public memory of the instrument.

Abstract red and gold composition

Calling these operations selections is not the same as calling them mistakes. An instrument without selection is barely an instrument. A telescope admits some wavelengths and excludes others. A map preserves some relations and suppresses others. A scientific interface must turn an impossible totality into something a person can inspect.

The difficulty is that each reduction can disappear behind the reality of its result. A neuron address looks like a stable object. An argmax winner looks like the most important participant. A minimal activating exemplar looks like the essence of an activation. A line between layer-wise winners looks like a path. A profile looks like a resident with a character. The more fluent the interface becomes, the easier it is to forget the operations that produced what we are seeing.

The project has already encountered several versions of this. L4-N923 appeared to be an extraordinary landmark across all eighty selected 541 routes, then won across eighty unrelated controls as well. The result did not reveal a 541-specific precursor; it revealed how a generic winner could dominate the chosen readout. Threads accurately recorded successive layer-wise argmax results, but their continuous lines supplied a visual claim of connected travel that the measurement had never made. Both episodes are documented elsewhere because they changed the instrument. Here they matter as examples of the same pressure: selection does not merely omit information. It can give the retained information a stronger ontology than it has earned.

The review's discussion of superposition and polysemanticity supplies another pressure on the address itself. Individual MLP neurons are not meaningless coordinates: the network gives them a privileged basis, and per-neuron behaviour can be empirically useful. But neurons are also often polysemantic. More features may be represented than there are neurons available to hold them separately, leaving features distributed across overlapping directions.

Mesocosm can therefore make a neuron spatially addressable without making it conceptually singular. A destination such as L5-N541 is a reproducible location under a declared function. It is not automatically one feature, one meaning or one mechanism. The address is real. The unity suggested by the address remains an empirical question.

Selection is not a flaw we can remove. It is a condition we have to keep visible.

This changes how the review's discussion of interpretability illusions should be applied here. The warning is not that an activation observation becomes worthless until someone performs an ablation. It is that the observation must not silently climb into a stronger kind of claim.

What the current evidence permits
Observation

This neuron won here under this model, corpus, tokenization and measurement rule.

Pattern

Across the available examples, this bounded regularity recurs often enough to describe and challenge.

Intervention

Under a particular manipulation and control design, a predicted downstream quantity changed in the declared direction.

Mechanism

A larger account connects the relevant features and computations, survives alternatives and explains the behaviour causally.

These are not stages through which every neuron is obliged to pass. Sometimes six observations support only a clue. Sometimes a pattern can be characterized without any obvious intervention that would isolate it. Sometimes an ablation produces a difference while leaving the complete mechanism unresolved. L5-N541 progressed further because its observed behaviour supported sharper questions, not because every address contains the same sequence waiting to be unlocked.

This is where the streetlight becomes more difficult than publication bias. Causal work at neuron-level resolution is often necessarily bespoke. A response to connected parallel members, an orthographic neighbourhood around whelm, a neuron that peaks on example, and a beginning-of-sequence winner do not naturally share one counterfactual dataset or one diagnostic intervention. The relevant position, replacement, control, metric and expected direction depend on the phenomenon.

The V12 work on 541 combined local ablation, replacement, injection, positional controls and matched-neuron controls because those operations were shaped around pair-completion behaviour. Even then, injection was dominated by generic disruption, and necessity, sufficiency and generation were not established. The intervention narrowed the interpretation without converting the neuron into a solved mechanism.

This is why I remain reluctant to place a generic laboratory inside the game. A standard suite could suggest that every observation can be investigated in the same way. The literature does not really offer such a universal recipe either. Activation patching depends on hand-built clean and counterfactual inputs, choices about direction, positions and metrics, and a prior understanding of the behaviour one is trying to isolate.

But necessary specificity has a selection effect of its own. Cases with clean contrasts and tractable outcomes are more likely to receive causal follow-up. Sparse, mixed or poorly localized observations may remain observational not because they are unimportant, but because the right experiment is unclear or prohibitively expensive. The strongest part of the archive will therefore be biased toward phenomena that were already unusually capable of becoming experiments.

The second distinction

The experiment may need to remain local to the phenomenon. The record of how that phenomenon was selected does not.

Abstract red and gold composition

The site is not outside the instrument. It is where scattered encounters become the durable public shape of what the instrument appears to find.

Inside Mesocosm, a player can land on a neuron, alter a string, watch a result survive or collapse, and continue elsewhere. On the site, those temporary movements become named records. A profile accumulates evidence. A writeup supplies sequence and interpretation. An image gives one result visual authority. Links make some residents central enough that later pages can assume them as shared history.

This is useful. It is also another transformation. The archive compresses the activity of the project into a collection a reader can traverse. Like every other instrument discussed here, it selects.

The answer cannot be to give every neuron equal treatment. Most have not been encountered under the current protocol. Many that have appeared support only thin observations. A demand for one complete profile per address would produce thousands of ceremonial placeholders, disguise differences in evidence quality, and consume the project in the name of avoiding selection.

Nor is random selection intrinsically more virtuous than curated depth. Random expeditions can estimate something about ordinary yield only when the sampling rule, substitutions, stopping criteria and denominator are designed for that purpose. A handful of Reddit requests is a reality stress test, not a prevalence study. Meanwhile, deliberately following a rich case remains the only way to learn what sustained investigation can sometimes achieve.

The important distinction is not curated versus random as good versus bad. It is knowing which question each mode can answer.

Mode 01 // Exemplar

How far can this go?

A case is selected because it supports depth. It can demonstrate a method, expose a phenomenon and sustain bespoke controls.

541 belongs here
Mode 02 // Expedition

What did we find here?

A destination is chosen before its promise is known. Sparse, mixed and inconclusive returns remain legitimate outcomes.

The N2222 request belongs here
Mode 03 // Survey

How often does this happen?

A sampling design estimates a distribution. It needs a declared population, denominator, selection rule and consistent stopping conditions.

Mesocosm has not done this yet

These modes can coexist, but they should not impersonate one another. L5-N541 can be an exemplary investigation without estimating the interpretability of arbitrary neurons. N2223 can preserve an unscreened encounter without becoming evidence for the typical failure rate of whittling. A future survey could sample addresses systematically, but it would sacrifice some of the freedom that allowed 541's investigation to evolve around the phenomenon.

This returns us to the split between what can be generalized and what must remain local. The experiment often belongs to the phenomenon. Pair completion called for one collection of interventions; an apparent orthographic neighbourhood might call for another; a positional landmark might be tested through carrier and position controls before semantic counterfactuals make sense at all.

The common structure can live in the record around those experiments.

Selection provenance
Was the neuron curated after prior exploration, requested externally, sampled by a declared rule, or substituted because the requested address was absent?
Measurement boundary
Which model, layer, corpus, tokenization, winner criterion and aggregation rule made this observation available?
Available denominator
How many corpus hits, routes, controls or sampled addresses sit behind the visible examples?
Resistance
Which outliers, failed controls, alternative clusters or generic effects prevented a cleaner description?
Evidence state
Is this a single observation, a recurring pattern, a bounded characterization, an intervention result or part of a larger mechanistic account?
Stopping point
Did the investigation stop because the evidence was sufficient, the sample was too sparse, the hypothesis failed, the experiment was unclear, or attention moved elsewhere?

None of this requires placing a universal laboratory inside the game. It asks for something less dramatic and perhaps more durable: a shared grammar for provenance and uncertainty. Session records can preserve the string, address, transformation and result. Atlas entries can distinguish how a star entered the record. The site can label an exemplar as an exemplar and an expedition as an expedition. Failed or sparse returns can remain visible without being inflated into feature-length articles.

This also keeps the game's purpose in proportion. Mesocosm does not need to complete the mechanistic interpretation of every star it makes visitable. Its instruments can remain invitations to form sharper questions: what survives this deletion, what changes across these layers, what does this control destroy, what would distinguish a positional response from a semantic one?

Sometimes the next honest operation will be another in-game transformation. Sometimes it will be a notebook, a script and a bespoke intervention outside the game. Sometimes it will be a record saying that six examples were not enough.

The starmap does not owe every star a secret. It owes the investigator an honest account of how a star became visible.

The brightest cases can remain bright. L5-N541 has earned its profile through the amount of evidence gathered around it, the controls it survived, the claims it lost and the questions it continued to support. The correction is not to dim that work until every neuron looks equal. It is to preserve the surrounding field well enough that one articulate resident does not become an accidental theory of all the others.

If the game makes someone curious enough to start writing tests, the design may already have won. The next responsibility is to remember why that test was written for this star, what selected it from the field, and what remained unresolved beyond the reach of its light.

Final return // Exemplary is not representative Keep the bright case. Keep the field.

An honest archive can hold deep investigations, arbitrary expeditions and eventual surveys without asking one kind of record to stand in for the others.