MESOCOSM //
Research writeup
GPT-2 Small // L5 → L6
Field return // adjacent-layer trajectories

When the threads branch.

A visual split in a game became a question about language, timing, and causality. Some paths dissolved under inspection. Others became more interesting.

The starmap resembles a routing diagram. That resemblance is the hypothesis being tested—not the conclusion. Its branches are repeated argmax trajectories, not demonstrated circuit edges.
Godot starmap showing threads from L5-N541 separating toward multiple layer-six destinations
01 // What the game saw

A split is a clue,
not a conclusion.

The same eighty reduced Moby-Dick text fragments were measured at every layer. At each layer, the “winner” is simply the MLP neuron with the largest activation observed anywhere across the fragment's token positions; a thread joins those successive winners. All eighty land on N541 at layer five, then repeatedly separate toward several recurring layer-six winners, producing visible branches. The picture gave us a reason to ask whether those branches reflect shared linguistic structure, temporal coincidence, or causal influence. It could not answer the question itself.

01

See

Recurring thread destinations appear as branches in the Godot starmap.

02

Count

Full sentences and reduced strings are grouped by their natural L6 winners.

03

Locate

The peak token reveals whether L5 and L6 events occur together or elsewhere.

04

Intervene

N541 is zeroed at its natural peak and each L6 response is measured again.

05

Challenge

Predictions are frozen, then tested on new sentences and paired disruptions.

02 // The first return

Not one successor.
Different effects.

The candidate branches did not respond uniformly when N541 was removed. Some fell at the same construction-completion token. One rose. Others had already peaked earlier and could not have been caused by the later N541 event. If removing N541 at its peak changes an L6 neuron at that same token, that supports a local causal contribution rather than sentence-level association alone.

Layer 5 // sourceN541construction completion

→ N2712

Successive / distributed recurrence

one by one · wave after wave · night after night

ABLATION // DECREASE

→ N1506

Intermittent frequency

now and then; a later explicit often can supply the signal independently.

ABLATION // DECREASE

→ N3051

Cyclic / continuing recurrence

round and round · again and again · over and over

ABLATION // DECREASE

↗ N2662

Direct alignment / correspondence

shoulder to shoulder and related X-to-X arrangements.

ABLATION // INCREASE
03 // Peak tracing

Timing separates
path from company.

A sentence-wide winner may peak on completely different words from N541. Tracing positions changed the picture: local continuations aligned on the same token; many apparent branches were simply elsewhere in the same long sentence.

A local continuation

one by one

N541 and N2712 peak on the second one. Removing N541 lowers N2712 by roughly half an activation unit and changes the destination.

An independent later cue

here and there … often

N541 peaks on there; N1506 peaks thirteen tokens later on often. Removing the earlier event barely changes the later explicit frequency signal.

An earlier maximum

up and down

N2374 peaks on and, before N541 peaks on down. A later intervention cannot rewrite an activation already computed. N2374 remains unchanged.

An opposite effect

shoulder to shoulder

N541 and N2662 share the completion token, but N2662 rises when N541 is removed. The relationship looks competitive rather than supportive.

80Full natural routes

Every sentence reproduced its recorded L5 and L6 destination before intervention.

14Same-token peaks

N541 and the natural L6 winner reached their maxima on the same token.

11Winner changes

Ten occurred in the same-token group; one came from N2662 rising past an unaffected winner.

0N2374 local effect

All four N2374 maxima occurred before the later N541 peak.

04 // Held-out panel

The neat story
fails usefully.

Before scoring, twenty-four new positive sentences and twenty-four paired disruptions were assigned to four predicted branches. The causal signs generalized better than the semantic labels.

Branch
Prediction
Beats control
Candidate win
L6 win
Ablation sign
2712
Successive recurrence
5 / 6
1 / 6
1 / 6
6 / 6 ↓
1506
Intermittent frequency
6 / 6
2 / 6
1 / 6
6 / 6 ↓
3051
Cyclic recurrence
6 / 6
5 / 6
5 / 6
6 / 6 ↓
2662
Direct correspondence
6 / 6
5 / 6
2 / 6
6 / 6 ↑
THE IMPORTANT CORRECTION // The predicted neuron won the four-candidate comparison in 13/24 cases, not 24/24. N2712 was not a general X by X winner; N1506 did not own every frequency-like idiom. The evidence supports graded, overlapping effects—not a routing table of semantic labels. N3051's disruption controls also retained some N541-mediated effect, weakening its construction specificity under this panel.
Repeated visual trajectories can identify candidate structure. They do not become semantic classes or causal edges merely because we draw lines between them.

The visual branching was doing something valuable: compressing repeated measurements into a form that made regularity noticeable. It surfaced a question that survived direct intervention.

But the experiment changed the language we should use. Some branches are same-token, differently signed relationships worth following. Others are earlier or unrelated maxima sharing a sentence. Even the strongest held-out result is graded and context-sensitive.

A fair description is: N541-origin strings separate into recurring L6 destinations whose timing, linguistic composition, and response to intervention differ.

05 // Inspection and reproduction

Open the
apparatus.

The bundle contains the exact three scripts used here, the reduced-string source packet, four CSV result tables, and a README documenting the pinned model and measurement regime. Full-corpus destination exports are identified but not duplicated in the compact bundle.