MESOCOSM
Paid work
Remote // Australia
Current paid priorities // 2026

Before we scale.

Mesocosm needs to establish what its instruments can honestly show before turning a one-book prototype into a larger system. Scientific judgment comes first. Engineering follows from it.

Priority 01 // seeking now

Mechanistic
interpretability.

Mesocosm is looking for a paid researcher or research engineer to examine an interactive visualisation built around GPT-2 Small. The first engagement is a technical and scientific review—not a request to endorse the project, and not necessarily a long-term commitment.

A useful first review could cover

01 // MEASURESWhat the current prototype extracts, transforms, and preserves.
02 // CLAIMSWhich interpretations are supported, uncertain, misleading, or still testable.
03 // CONTROLSAppropriate baselines, comparisons, interventions, and validation experiments.
04 // PEDAGOGYLanguage and interaction choices that invite curiosity without false certainty.
05 // GENERALISATIONWhat should remain stable across different texts, datasets, layers, or models.
06 // ROADMAPA small, credible research programme for a public demo and its supporting records.
Possible first arrangement

Paid consultation

A walkthrough of the current build and analysis pipeline, followed by written or recorded technical feedback and a discussion of next experiments.

Who may fit

Researcher or research engineer

Someone familiar with activation analysis, representation research, experimental design, reproducible ML tooling, or adjacent interpretability practice. Critical perspectives are welcome.

The point is not to make the project sound more scientific. It is to identify which parts deserve confidence, which parts need stronger tests, and which parts should be described more carefully or discarded.
A group of people examining a vast open black-box structure illuminated in red
Shared examination // the black box as a place of work
Priority 02 // being scoped

Beyond a
single book.

The current prototype works with Moby-Dick and precomputed caches. A clear early engineering win would be to make the workflow reproducible across several books or other text collections—once the scientific review has helped establish what is worth computing.

A colossal black-box architecture glowing with red vertical data structures
From singular object to durable architecture

A first milestone might

01 // AUDITDocument the existing Moby-Dick preprocessing and cache workflow.
02 // INGESTAccept and consistently tokenise multiple source texts.
03 // COMPUTERun inference and precompute the reviewed activation-derived material.
04 // PROVENANCERecord model, tokenizer, layer, source, code version, and analysis settings.
05 // STOREUse a versioned, inspectable format that can survive changes in the prototype.
06 // SERVELet Godot load or stream selected material without holding the whole collection in memory.
Likely shape

Defined engineering milestone

A documented pipeline, a representative multi-book dataset, a stable handoff to Godot, and measurements of storage and compute requirements.

Likely fit

Python / ML systems engineer

Someone comfortable with model inference, data formats, reproducibility, performance constraints, and designing an interface between research code and an interactive application.

This is a sketch, not a frozen specification. The exact material to cache—and therefore the architecture used to hold it—should follow the interpretability review rather than precede it.
Open inquiry

Begin with
a conversation.

If your work sits near either problem, you are welcome to get in touch. A first message can be brief: who you are, what you noticed here, relevant work or publications, your availability, and your usual rate or preferred way of scoping consultation.