Automating Interpretability Research: First Steps
As frontier AI research increasingly becomes automated, our tools for understanding models need to keep up.
In this post, we share our work on automating interpretability research and promising first results from our system.
Our System
Our system consists of an idea generator agent, an experimental plan agent, and a research execution agent.
Idea Generator
Our idea generation agent is simple: it coordinates an idea proposer, a quality reviewer, and a novelty reviewer. The generator comes up with its ideas from scratch. It can optionally take in a prompt specifying a research area to focus on. The best prompt we've tried using is a post by Jack Lindsey listing important open problems in interpretability.
As models become stronger at following instructions, we’ve found it far easier to specify the actions that should occur in the prompt rather than hardcoding them into an agentic framework. As such, we’re also providing a diagram of the intended path the agents should take. This is the path outlined for the idea generation agent to take as specified by its prompt.
The generator call is a call to Opus-5.5 within the Claude Code harness, and the quality reviewer and novelty reviewers are both calls to GPT-6 Astra in the Codex harness. The quality reviewer does not have internet access; the novelty reviewer does. The generator and reviewer agents continue their previous sessions during each cycle of review.
We benchmarked the quality reviewer by selecting existing conference papers, generating ideas out of them, and labelling them good/bad and measuring agreement between the quality reviewer and our labels. We ended up with a small benchmark of 17 ideas; our reviewer rejected every bad idea and accepted 5 of 6 strong ideas.
Sample Ideas Generated
Experimental Plan
The experimental plan agent also follows a relatively simple pattern of a generator and reviewer agent. We use a Codex agent with GPT-6 Astra for both generator and reviewer.
GPT-6 Astra and Opus 5.5 were surprisingly poor at creating research plans that would validate a hypothesis; when called without our prompt, they often propose datasets that cannot reasonably teach our desired behavior, highly inefficient steps, never base code off of existing projects, and more. Our prompt instructs Astra to read the closest reputable paper to the proposed idea and adapt its research plan to our idea. This is extremely effective at mitigating these poor planning errors, but we expect to change this in the future as this inherently limits the creativity of the experimental plan.
Research Execution Agent
Our research execution agent has become a single GPT-6 Astra Codex agent with high reasoning that follows the following plan:
The biggest issues we’ve encountered with the research agent have been: experiments that use excessive time/resources to scale an idea without validating it, inefficient pipelines, most notably, repeatedly using Transformer's .generate() command to inference a model instead of vLLM, and difficulty adapting the research plan. As we develop this pipeline, we have manual intervention whenever the system hiccups to keep it on track.
Stacking Research Advancements
This writeup has focused on a single serial interpretability research system. After an advancement is made, importantly, the work needs to be built off of and expanded upon by other agents. We’ve experimented doing so by building off of the surface area decomposition and quality diversity databases in our previous writeup, Scaling Autonomous Research to Thousands of Agents. However, until we speed up our current interpretability agent’s iteration time, we’ve decided to prioritize work on the single agent.
Findings
Full-Context NLA
Natural language autoencoders take in an activation vector, translate it into a natural language readout, and reconstruct the original vector from the readout. Typically, these methods only work on a single token of context. Our autonomous interpretability research system extended the NLA to a full-context NLA, where it takes in and reconstructs an activation vector from every token in the context.
Architecture
More specifically, the primary change was adding an extra activation per token in the prompt and adding more placeholder slots for the reconstructor to fill.
Loss Functions
Ordinary NLA trains with a supervised fine-tuning stage and an RL stage. The SFT stage creates summaries of the readable text and trains the reader to output these summaries, while the reconstructor is trained to turn them back into activations. This creates a strong start before the RL phase, which optimizes for reconstruction.
Concretely:
SFT Loss function. For an activation and a summary of the text up to that token, the reader is trained with cross-entropy on the summary and the reconstructor with squared error:
RL Loss function. The reader samples a description and is rewarded for reconstruction, with a KL penalty toward its initialization; the reconstructor keeps regressing onto from the sampled descriptions:
These are the original paper’s objectives, with activations normalized to unit length. For optimization, the reader is updated with GRPO over a group of sampled descriptions per activation, and the reader and reconstructor updates are taken in parallel on each batch.
The Full Context NLA system had an identical setup: an SFT stage that trained on summaries of data and an RL phase to optimize for reconstruction. Two RL loss functions were tested: one that equally weighted the reconstruction loss from every token in the context, and one that weighted the loss of every token by how much it was attended to by the final token in the sequence.
For a context of tokens with activations and reconstructions read from the placeholder slots, all normalized to unit length:
Unweighted loss function.
Weighted loss function.
Here is the attention that head of the original model’s next layer pays from the final token to position , normalized over positions. It is computed from the original activations and kept frozen; that layer attends to at most the last 1,024 positions, so earlier positions get zero weight. In our full-context RL runs, the reader’s reward was the negative error, (no logarithm), with the same KL penalty, and the reconstructor regressed on .
Results
We trained for 10% of the SFT data and 3% of the RL data specified in the original NLA paper. We used this recipe for three models: a reproduction of the NLA setup, a full context model trained on the unweighted loss function, and a full context model trained on the weighted loss function. All models are trained on Gemma-3 12B. Each RL step takes considerably longer than the original pipeline. We also changed the summarization model to generate the SFT data from Sonnet 4.6 to GPT 5.6 Luna. All other hyperparameters remain constant, except for the learning rate on the unweighted RL run, which was reduced from 1e-5 to 3e-6 after observing a blowup at 1e-5.
Variance explained (%)
Last token
Average over tokens
Attention-weighted tokens
Classification
Source:
Asked: About Sports? (expected YES)
User attributes
Source:
Asked: Judge given “tour guide”: does the readout support it?
Safety attribution
Source:
Model answer: B → A with the harmful request.
Asked: Does the readout attribute the answer to safety? (a mention, not a verified cause)
Discussion
Qualitatively, the readouts appear richer and notice new characteristics; in the example we give, full context readouts make an inference about the identity of who’s described far more often, and occasionally flag safety attempts. However, the time for an RL update rose from 2.6 minutes to 4.6 minutes for a modest improvement. In our testing, the full context NLA was also very compute efficient, getting strong performance after 10% of the SFT and 3% of the RL budget of the original paper. There are several issues underlying NLAs mentioned in the original paper that this approach shares, such as how the objective is not guaranteed to produce faithful readouts, and the difficulty in evaluating these methods. Further work is needed on addressing these issues before we can see the true usefulness of the full context NLA.
A Study of Monitor Placement
Layer and token-position choice for Jacobian-lens and linear-probe safety monitors in three small open models.Note: The following two writeups were generated by another agent that produces experimental writeups.
Summary
If you monitor a language model by reading its activations, you have to choose a layer and a token position to read from. We tested how much that choice matters for two monitors. The first is a training-free readout through the Jacobian lens ("JSpace"): a fixed comparison of the words harmful, unsafe against safe, benign. The second is a supervised linear probe. Both were run on Llama‑3.2‑3B, Qwen3‑4B and Gemma‑3‑4B‑it, for three tasks on supplied text: is the prompt harmful, is the response harmful, and does the response refuse.
What we found
- Placement matters a lot for the JSpace readout, and the best placement depends on the model and task. Across all layer and position choices, AUROC on untouched data ranges from 0.28 to 0.79. Within each model–task cell, the best placement beats the worst by 0.25–0.40. Among the prompt-harm peaks, Gemma's is at layer 8 (of 0–33) at the end of its chat template, Qwen's at layer 25 (of 0–35) on the final content token, and Llama's at layers 12–13 (of 0–27) on the last word. No single layer or position rule fits all nine model–task cells.
- A placement chosen on development data carried over to new data from the same source. In 6 of 9 cells, the frozen placement was exactly the best of all 54–99 placements on 2,400 untouched prompt groups. The largest shortfall from the best placement in hindsight was 0.016 AUROC.
- Even the best-placed JSpace readout is a weak monitor. Frozen JSpace AUROCs are 0.575–0.792, against 0.979–0.999 for the probes. At the equal-cost threshold, JSpace misclassifies 16–39 percentage points more examples than the probe. All nine cells fail the pre-registered usefulness test. On 450 human-written XSTest prompts, the same frozen policies wrongly flag 92–108 of 250 safe prompts.
- On the saved probe placements, the probe was much less sensitive to placement. Its layer‑0 probes scored 0.88–0.95. Every other saved probe placement scored 0.969–0.999. This study did not save a full per-layer probe map on the confirmation data.
So the answer to "where should a monitor read?" is narrower than a placement manual. For this fixed-word readout, find the placement for each model and task on a few thousand labeled development examples. That search worked here. But the readout stayed far below a trained probe wherever it read.
Introduction
The Jacobian lens (Gurnee et al., Verbalizable Representations Form a Global Workspace in Language Models) asks what an intermediate activation is disposed to make the model say. For each layer ℓ, it averages, over 1,000 ordinary text contexts, the Jacobian of the final-layer residual stream with respect to the layer-ℓ residual. The result is one matrix, J_ℓ. Applying J_ℓ and then the model's own unembedding to an activation gives a score for every vocabulary token. The logit lens is the special case J_ℓ = I. The paper reports that post-trained models show "safety concerns" in this space while still reading a user's message. It also argues that the lens "could be highly useful for alignment monitoring", while declining to claim that monitoring the J‑space is sufficient.
That suggests a cheap monitor with no task training. Read the activation at some layer and token, project it through the lens, and compare how strongly it points toward harmful/unsafe versus safe/benign. If it works, the open question is where to read. The paper finds its workspace in a middle-to-late layer range of Claude models, and the right token in a chat transcript is not obvious.
There is an obvious simpler explanation for any success: the readout may detect topical words ("kill", "abuse") rather than harmful intent. A trained linear probe (Oldfield et al., Beyond Linear Probes, released linear baseline) is a strong supervised alternative. It needs labels; the lens readout does not.
What would establish the benefit? A placement fixed in advance on development data, then tested on untouched data, must (1) lose at most 0.02 AUROC against a placement re-tuned on fresh data, (2) have error cost within 0.02 of the probe at three false-negative:false-positive cost ratios, and (3) beat always-flag/never-flag decisions. These criteria were fixed before confirmation in a frozen protocol.
Architecture
Results
Three patterns, each limited to these models and labels:
- The best layer is not a fixed depth. The prompt-harm peaks sit at relative depths of about 0.44–0.48 (Llama, layers 12–13 of 0–27), 0.71 (Qwen, 25 of 0–35) and 0.24 (Gemma, 8 of 0–33).
- Position can matter as much as layer. For Qwen refusal, the final content token rises to 0.76 while the last lexical token stays near or below chance (0.35–0.57). For Gemma, the template end gives the best prompt-harm (0.79, layer 8) and response-harm (0.74, layer 5) readouts, at early layers where the other two positions score only 0.48–0.62.
- Refusal is the weakest task. The refusal vocabulary scores below 0.5 at 37–49 of the 54–99 placements, so many placements rank refusals backwards. The logit lens beat JSpace on all three refusal tasks (0.74, 0.81 and 0.78 against 0.58, 0.76 and 0.66).
The placement found on development data was, in most cells, the best one on the confirmation set:
| Model · task | Frozen placement (layer: position) | Its AUROC | Best of all placements on confirmation | Gap |
|---|---|---|---|---|
| Llama · prompt harm | 13:last-lexical | 0.669 | 12:last-lexical, 0.676 | 0.007 |
| Llama · response harm | 25:final-content | 0.660 | same | 0 |
| Llama · refusal † | 18:final-content | 0.575 | same | 0 |
| Qwen · prompt harm | 25:final-content | 0.761 | same | 0 |
| Qwen · response harm | 25:final-content | 0.775 | same | 0 |
| Qwen · refusal | 32:final-content | 0.756 | 26:final-content, 0.757 | 0.0002 |
| Gemma · prompt harm | 8:serialized-end | 0.792 | same | 0 |
| Gemma · response harm | 5:serialized-end | 0.739 | same | 0 |
| Gemma · refusal † | 19:final-content | 0.656 | 14:final-content, 0.672 | 0.016 |
The "best in hindsight" column is picked on the evaluation data itself, so it flatters the comparison; the small gaps are therefore conservative. This is our descriptive recomputation from the saved scores. The pre-registered regret test is narrower. In seven cells, the source and target searches picked the same placement, so regret is zero by identity; that is not seven independent demonstrations of transfer. The two distinct pairs have regret −0.018 (Llama prompt harm, adjusted interval [−0.057, 0.021]) and 0.0002 (Qwen refusal, [−0.043, 0.044]). Both intervals cross the 0.02 margin, so both are unresolved. All of this is same-source transfer between partitions of one dataset. The earlier plan to hold out a whole category was dropped, so this study makes no category-transfer claim.
Queryable NLA: Negative Result
Summary
Could we ask an activation what it contains, rather than accept whichever description a natural language autoencoder chooses to produce? We tried a simple interface: fine-tune an NLA to describe activations as subject–relation–value tables, then supply the subject and relation and let it fill the value. The tables were a means to make the reader queryable. Familiar forced fields preserved reconstruction, but new queries were unreliable.
Architecture
Starting from the released Qwen2.5-7B layer-20 NLA, we fine-tuned its reader on table conversions of its own descriptions and adapted its reconstructor to recover activations from those tables. Constrained generation enforced three fields per row and at least six rows. To query it, we fixed the first row’s subject and relation; the model generated the value and remaining rows. It received an activation, not the source text or an answer key.
Results
On 128 development examples, fixing a subject and relation already present in an earlier readout barely changed reconstruction: 60.24% → 60.12% variance explained. This checked whether forcing familiar fields preserved reconstruction, not whether the reader could answer a new question.
We then assessed whether the outputs actually contained requested answers, including answers appearing outside the first row.
| Query test | Released NLA, prose prefix | Table-adapted NLA |
|---|---|---|
| 192 synthetic contextual queries | 1/192 | 5/192 |
| 8 factual queries from natural passages | Not tested | 0/8 |
The 192 queries covered locations, reported beliefs, and activities across 24 source templates. Manual scoring found only isolated correct answers; ambiguous cases give ranges of 1–2/192 and 3–5/192. In one credited table answer, replacing the activation with a conflicting example left the answer unchanged. These results do not establish an advantage over ordinary NLA.
The eight natural-passage queries were selected for explicit, checkable facts. All outputs had valid table structure, but none answered correctly—even elsewhere in the table. Reconstruction on this smaller set was 47.98% unforced versus 43.30% queried, using its own variance denominator. For example:
| Forced subject and relation | Answer in source | Generated value, verbatim |
|---|---|---|
| InScribe Module 2 — costs | £25 | free online course |
| breakfast cereal — served with | milk | historical and scientific description of rubber |
| Syv Systre — country | Norway | United States |
Preserving reconstruction did not make the reader reliably answer queries.
Discussion
The current autonomous interpretability research system still has room for improvement. For resource-intensive experiments such as the full-context NLA and queryable NLA, we still had occasional manual interventions, and our system has yet to generate a paradigm shift in interpretability research. We’re working on making this system stronger until we have boundless open-ended interpretability research able to keep up with capabilities.
