Automating Interpretability Research: First Steps

As frontier AI research increasingly becomes automated, our tools for understanding models need to keep up.

In this post, we share our work on automating interpretability research and promising first results from our system.

Our System

Our system consists of an idea generator agent, an experimental plan agent, and a research execution agent.

Idea Generator

Our idea generation agent is simple: it coordinates an idea proposer, a quality reviewer, and a novelty reviewer. The generator comes up with its ideas from scratch. It can optionally take in a prompt specifying a research area to focus on. The best prompt we've tried using is a post by Jack Lindsey listing important open problems in interpretability.

Research advice feeds a generator, whose ideas go to a quality reviewer and a novelty reviewer; feedback returns to the generator, and accepted ideas advance.
A diagram showing the layout of the idea generation agent.

As models become stronger at following instructions, we’ve found it far easier to specify the actions that should occur in the prompt rather than hardcoding them into an agentic framework. As such, we’re also providing a diagram of the intended path the agents should take. This is the path outlined for the idea generation agent to take as specified by its prompt.

Research advice, read papers, find a limitation, simple method, propose a test, rank ideas
The sequence of events outlined in the idea generation agent’s prompt.

The generator call is a call to Opus-5.5 within the Claude Code harness, and the quality reviewer and novelty reviewers are both calls to GPT-6 Astra in the Codex harness. The quality reviewer does not have internet access; the novelty reviewer does. The generator and reviewer agents continue their previous sessions during each cycle of review.

We benchmarked the quality reviewer by selecting existing conference papers, generating ideas out of them, and labelling them good/bad and measuring agreement between the quality reviewer and our labels. We ended up with a small benchmark of 17 ideas; our reviewer rejected every bad idea and accepted 5 of 6 strong ideas.

Sample Ideas Generated

Experimental Plan

A research idea goes to a plan generator (Astra medium), whose experimental plan goes to a plan reviewer (Astra medium); the reviewer either sends revisions back or advances the plan.
The experimental plan generation agent’s prompt.
Research idea, define claims, reuse baseline, minimal changes, design tests, budget and check
The sequence of events outlined in the experimental plan generation agent’s prompt.

The experimental plan agent also follows a relatively simple pattern of a generator and reviewer agent. We use a Codex agent with GPT-6 Astra for both generator and reviewer.

GPT-6 Astra and Opus 5.5 were surprisingly poor at creating research plans that would validate a hypothesis; when called without our prompt, they often propose datasets that cannot reasonably teach our desired behavior, highly inefficient steps, never base code off of existing projects, and more. Our prompt instructs Astra to read the closest reputable paper to the proposed idea and adapt its research plan to our idea. This is extremely effective at mitigating these poor planning errors, but we expect to change this in the future as this inherently limits the creativity of the experimental plan.

Research Execution Agent

Our research execution agent has become a single GPT-6 Astra Codex agent with high reasoning that follows the following plan:

Accepted plan, pilot, experiment, evaluate, write-up, with evaluation able to send revisions back to the pilot
The sequence of events outlined in the research execution agent’s prompt.

The biggest issues we’ve encountered with the research agent have been: experiments that use excessive time/resources to scale an idea without validating it, inefficient pipelines, most notably, repeatedly using Transformer's .generate() command to inference a model instead of vLLM, and difficulty adapting the research plan. As we develop this pipeline, we have manual intervention whenever the system hiccups to keep it on track.

Stacking Research Advancements

This writeup has focused on a single serial interpretability research system. After an advancement is made, importantly, the work needs to be built off of and expanded upon by other agents. We’ve experimented doing so by building off of the surface area decomposition and quality diversity databases in our previous writeup, Scaling Autonomous Research to Thousands of Agents. However, until we speed up our current interpretability agent’s iteration time, we’ve decided to prioritize work on the single agent.

Findings

Full-Context NLA

Natural language autoencoders take in an activation vector, translate it into a natural language readout, and reconstruct the original vector from the readout. Typically, these methods only work on a single token of context. Our autonomous interpretability research system extended the NLA to a full-context NLA, where it takes in and reconstructs an activation vector from every token in the context.

Architecture

Ordinary NLA: one activation goes through the reader to readable text of at most 150 tokens, then through the reconstructor to one reconstructed activation. Full-context NLA: activations for every token go through the same path and every activation is reconstructed.
A general outline of the change made by the full context NLA

More specifically, the primary change was adding an extra activation per token in the prompt and adding more placeholder slots for the reconstructor to fill.

Single token: a prompt and one activation embedding go to the reader; its text and one placeholder slot go to the reconstructor, whose head outputs one reconstructed activation. Full context: one activation embedding per token, one placeholder slot per token, and the same head outputs an activation at every slot.
A more concrete outline of the change made by the NLA

Loss Functions

Ordinary NLA trains with a supervised fine-tuning stage and an RL stage. The SFT stage creates summaries of the readable text and trains the reader to output these summaries, while the reconstructor is trained to turn them back into activations. This creates a strong start before the RL phase, which optimizes for reconstruction.

Concretely:

SFT Loss function. For an activation hh and a summary ss of the text up to that token, the reader AVϕAV_\phi is trained with cross-entropy on the summary and the reconstructor ARθAR_\theta with squared error:

Lreader(ϕ)=− E(h,s)∑ilog⁡AVϕ ⁣(si∣s<i, h)\mathcal{L}_{\text{reader}}(\phi) = -\,\mathbb{E}_{(h,s)} \sum_{i} \log AV_\phi\!\left(s_i \mid s_{<i},\, h\right) Lreconstructor(θ)=E(h,s)∥h−ARθ(s)∥22\mathcal{L}_{\text{reconstructor}}(\theta) = \mathbb{E}_{(h,s)} \left\| h - AR_\theta(s) \right\|_2^2

RL Loss function. The reader samples a description zz and is rewarded for reconstruction, with a KL penalty toward its initialization; the reconstructor keeps regressing onto hh from the sampled descriptions:

max⁡ϕ  Eh Ez∼AVϕ(⋅∣h)[−log⁡∥h−ARθ(z)∥22]−β DKL ⁣(AVϕ ∥ AVϕinit)\max_\phi \; \mathbb{E}_{h}\, \mathbb{E}_{z \sim AV_\phi(\cdot \mid h)} \left[ -\log \left\| h - AR_\theta(z) \right\|_2^2 \right] - \beta\, D_{\mathrm{KL}}\!\left( AV_\phi \,\|\, AV_{\phi_{\text{init}}} \right) min⁡θ  Eh Ez∼AVϕ(⋅∣h)∥h−ARθ(z)∥22\min_\theta \; \mathbb{E}_{h}\, \mathbb{E}_{z \sim AV_\phi(\cdot \mid h)} \left\| h - AR_\theta(z) \right\|_2^2

These are the original paper’s objectives, with activations normalized to unit length. For optimization, the reader is updated with GRPO over a group of sampled descriptions per activation, and the reader and reconstructor updates are taken in parallel on each batch.

The Full Context NLA system had an identical setup: an SFT stage that trained on summaries of data and an RL phase to optimize for reconstruction. Two RL loss functions were tested: one that equally weighted the reconstruction loss from every token in the context, and one that weighted the loss of every token by how much it was attended to by the final token in the sequence.

For a context of TT tokens with activations h1,…,hTh_1, \dots, h_T and reconstructions h^1,…,h^T\hat h_1, \dots, \hat h_T read from the placeholder slots, all normalized to unit length:

Unweighted loss function.

Eunweighted=1T∑t=1T∥ht−h^t∥22E_{\text{unweighted}} = \frac{1}{T} \sum_{t=1}^{T} \left\| h_t - \hat h_t \right\|_2^2

Weighted loss function.

Eweighted=∑t=1Twt∥ht−h^t∥22,wt=1K∑k=1Kak(T→t)E_{\text{weighted}} = \sum_{t=1}^{T} w_t \left\| h_t - \hat h_t \right\|_2^2, \qquad w_t = \frac{1}{K} \sum_{k=1}^{K} a_k(T \to t)

Here ak(T→t)a_k(T \to t) is the attention that head kk of the original model’s next layer pays from the final token to position tt, normalized over positions. It is computed from the original activations and kept frozen; that layer attends to at most the last 1,024 positions, so earlier positions get zero weight. In our full-context RL runs, the reader’s reward was the negative error, r=−Er = -E (no logarithm), with the same KL penalty, and the reconstructor regressed on EE.

Results

We trained for 10% of the SFT data and 3% of the RL data specified in the original NLA paper. We used this recipe for three models: a reproduction of the NLA setup, a full context model trained on the unweighted loss function, and a full context model trained on the weighted loss function. All models are trained on Gemma-3 12B. Each RL step takes considerably longer than the original pipeline. We also changed the summarization model to generate the SFT data from Sonnet 4.6 to GPT 5.6 Luna. All other hyperparameters remain constant, except for the learning rate on the unweighted RL run, which was reduced from 1e-5 to 3e-6 after observing a blowup at 1e-5.

Ordinary NLAWeighted contextUnweighted context

Variance explained (%)

Last token

-200204060050100RL updatesOrdinary NLA, 0 updates: 19.4%Ordinary NLA, 50 updates: 35.2%Ordinary NLA, 100 updates: 48.6%Weighted context, 0 updates: 4.0%Weighted context, 50 updates: 31.5%Weighted context, 100 updates: 39.6%Unweighted context, 0 updates: 11.5%Unweighted context, 50 updates: 12.4%Unweighted context, 100 updates: 13.9%

Average over tokens

-200204060050100RL updatesWeighted context, 0 updates: -2.6%Weighted context, 50 updates: 29.6%Weighted context, 100 updates: 29.3%Unweighted context, 0 updates: 32.2%Unweighted context, 50 updates: 36.3%Unweighted context, 100 updates: 38.7%

Attention-weighted tokens

-200204060050100RL updatesWeighted context, 0 updates: 1.3%Weighted context, 50 updates: 35.4%Weighted context, 100 updates: 37.7%Unweighted context, 0 updates: -14.9%Unweighted context, 50 updates: 7.1%Unweighted context, 100 updates: 17.4%
Reconstruction variance explained at 0, 50, and 100 RL updates, after 100 SFT updates per component.

Classification

100%75%50%25%0%237/304234/304237/304234/304ReleasedNLAOurordinaryWeightedcontextUnweightedcontext

Source:

Asked: About Sports? (expected YES)

User attributes

100%75%50%25%0%0/320/3222/3225/32ReleasedNLAOurordinaryWeightedcontextUnweightedcontext

Source:

Asked: Judge given “tour guide”: does the readout support it?

Safety attribution

100%75%50%25%0%2/20/21/21/2ReleasedNLAOurordinaryWeightedcontextUnweightedcontext

Source:

Model answer: B → A with the harmful request.

Asked: Does the readout attribute the answer to safety? (a mention, not a verified cause)

Readouts from the released Gemma NLA and our three readers (ordinary, weighted, and unweighted full context, each after 100 RL updates) on the same target activations. Click a source to compare them.

Discussion

Qualitatively, the readouts appear richer and notice new characteristics; in the example we give, full context readouts make an inference about the identity of who’s described far more often, and occasionally flag safety attempts. However, the time for an RL update rose from 2.6 minutes to 4.6 minutes for a modest improvement. In our testing, the full context NLA was also very compute efficient, getting strong performance after 10% of the SFT and 3% of the RL budget of the original paper. There are several issues underlying NLAs mentioned in the original paper that this approach shares, such as how the objective is not guaranteed to produce faithful readouts, and the difficulty in evaluating these methods. Further work is needed on addressing these issues before we can see the true usefulness of the full context NLA.

A Study of Monitor Placement

Layer and token-position choice for Jacobian-lens and linear-probe safety monitors in three small open models.Note: The following two writeups were generated by another agent that produces experimental writeups.

Summary

If you monitor a language model by reading its activations, you have to choose a layer and a token position to read from. We tested how much that choice matters for two monitors. The first is a training-free readout through the Jacobian lens ("JSpace"): a fixed comparison of the words harmful, unsafe against safe, benign. The second is a supervised linear probe. Both were run on Llama‑3.2‑3B, Qwen3‑4B and Gemma‑3‑4B‑it, for three tasks on supplied text: is the prompt harmful, is the response harmful, and does the response refuse.

What we found

  • Placement matters a lot for the JSpace readout, and the best placement depends on the model and task. Across all layer and position choices, AUROC on untouched data ranges from 0.28 to 0.79. Within each model–task cell, the best placement beats the worst by 0.25–0.40. Among the prompt-harm peaks, Gemma's is at layer 8 (of 0–33) at the end of its chat template, Qwen's at layer 25 (of 0–35) on the final content token, and Llama's at layers 12–13 (of 0–27) on the last word. No single layer or position rule fits all nine model–task cells.
  • A placement chosen on development data carried over to new data from the same source. In 6 of 9 cells, the frozen placement was exactly the best of all 54–99 placements on 2,400 untouched prompt groups. The largest shortfall from the best placement in hindsight was 0.016 AUROC.
  • Even the best-placed JSpace readout is a weak monitor. Frozen JSpace AUROCs are 0.575–0.792, against 0.979–0.999 for the probes. At the equal-cost threshold, JSpace misclassifies 16–39 percentage points more examples than the probe. All nine cells fail the pre-registered usefulness test. On 450 human-written XSTest prompts, the same frozen policies wrongly flag 92–108 of 250 safe prompts.
  • On the saved probe placements, the probe was much less sensitive to placement. Its layer‑0 probes scored 0.88–0.95. Every other saved probe placement scored 0.969–0.999. This study did not save a full per-layer probe map on the confirmation data.

So the answer to "where should a monitor read?" is narrower than a placement manual. For this fixed-word readout, find the placement for each model and task on a few thousand labeled development examples. That search worked here. But the readout stayed far below a trained probe wherever it read.

Introduction

The Jacobian lens (Gurnee et al., Verbalizable Representations Form a Global Workspace in Language Models) asks what an intermediate activation is disposed to make the model say. For each layer ℓ, it averages, over 1,000 ordinary text contexts, the Jacobian of the final-layer residual stream with respect to the layer-ℓ residual. The result is one matrix, J_ℓ. Applying J_ℓ and then the model's own unembedding to an activation gives a score for every vocabulary token. The logit lens is the special case J_ℓ = I. The paper reports that post-trained models show "safety concerns" in this space while still reading a user's message. It also argues that the lens "could be highly useful for alignment monitoring", while declining to claim that monitoring the J‑space is sufficient.

That suggests a cheap monitor with no task training. Read the activation at some layer and token, project it through the lens, and compare how strongly it points toward harmful/unsafe versus safe/benign. If it works, the open question is where to read. The paper finds its workspace in a middle-to-late layer range of Claude models, and the right token in a chat transcript is not obvious.

There is an obvious simpler explanation for any success: the readout may detect topical words ("kill", "abuse") rather than harmful intent. A trained linear probe (Oldfield et al., Beyond Linear Probes, released linear baseline) is a strong supervised alternative. It needs labels; the lens readout does not.

What would establish the benefit? A placement fixed in advance on development data, then tested on untouched data, must (1) lose at most 0.02 AUROC against a placement re-tuned on fresh data, (2) have error cost within 0.02 of the probe at three false-negative:false-positive cost ratios, and (3) beat always-flag/never-flag decisions. These criteria were fixed before confirmation in a frozen protocol.

Architecture

Two monitors and three token positions
Schematic. (A) The three token positions compared, shown on a real XSTest prompt with an illustrative word-level split. The last lexical token is the last token containing a letter or digit. The final content token is the last non-whitespace token of the prompt (or of the response, for response tasks). The serialized end is the last token of the full model input. Llama and Qwen inputs carry no chat template, so their serialized end is the final content token on every input; only Gemma has three distinct positions. (B) The JSpace readout uses one token and no task training. The probe averages activations over the prompt or response span and is trained on 30k–69k labeled rows.

Results

AUROC of the JSpace readout at every source layer and token position on the untouched confirmation set. Each line is one token position. The open circle is the placement frozen from development data. The dashed line is the frozen probe, which uses a single placement and is not a per-layer sweep. † marks Llama and Gemma refusal: their readout failed its development check (p = 0.69 and 0.023 against a threshold of 0.0056), so those curves are descriptive only.

Three patterns, each limited to these models and labels:

  • The best layer is not a fixed depth. The prompt-harm peaks sit at relative depths of about 0.44–0.48 (Llama, layers 12–13 of 0–27), 0.71 (Qwen, 25 of 0–35) and 0.24 (Gemma, 8 of 0–33).
  • Position can matter as much as layer. For Qwen refusal, the final content token rises to 0.76 while the last lexical token stays near or below chance (0.35–0.57). For Gemma, the template end gives the best prompt-harm (0.79, layer 8) and response-harm (0.74, layer 5) readouts, at early layers where the other two positions score only 0.48–0.62.
  • Refusal is the weakest task. The refusal vocabulary scores below 0.5 at 37–49 of the 54–99 placements, so many placements rank refusals backwards. The logit lens beat JSpace on all three refusal tasks (0.74, 0.81 and 0.78 against 0.58, 0.76 and 0.66).

The placement found on development data was, in most cells, the best one on the confirmation set:

Model · taskFrozen placement (layer: position)Its AUROCBest of all placements on confirmationGap
Llama · prompt harm13:last-lexical0.66912:last-lexical, 0.6760.007
Llama · response harm25:final-content0.660same0
Llama · refusal †18:final-content0.575same0
Qwen · prompt harm25:final-content0.761same0
Qwen · response harm25:final-content0.775same0
Qwen · refusal32:final-content0.75626:final-content, 0.7570.0002
Gemma · prompt harm8:serialized-end0.792same0
Gemma · response harm5:serialized-end0.739same0
Gemma · refusal †19:final-content0.65614:final-content, 0.6720.016

The "best in hindsight" column is picked on the evaluation data itself, so it flatters the comparison; the small gaps are therefore conservative. This is our descriptive recomputation from the saved scores. The pre-registered regret test is narrower. In seven cells, the source and target searches picked the same placement, so regret is zero by identity; that is not seven independent demonstrations of transfer. The two distinct pairs have regret −0.018 (Llama prompt harm, adjusted interval [−0.057, 0.021]) and 0.0002 (Qwen refusal, [−0.043, 0.044]). Both intervals cross the 0.02 margin, so both are unresolved. All of this is same-source transfer between partitions of one dataset. The earlier plan to hold out a whole category was dropped, so this study makes no category-transfer claim.

Frozen policies versus probes
AUROC of the frozen JSpace placement, the development-selected logit-lens placement, and the frozen linear probe. The top block is the WildGuard-mirror confirmation set; the bottom block is XSTest, scored with the same frozen policies and no re-tuning. The probe gap is consistent across models and tasks.

Queryable NLA: Negative Result

Summary

Could we ask an activation what it contains, rather than accept whichever description a natural language autoencoder chooses to produce? We tried a simple interface: fine-tune an NLA to describe activations as subject–relation–value tables, then supply the subject and relation and let it fill the value. The tables were a means to make the reader queryable. Familiar forced fields preserved reconstruction, but new queries were unreliable.

Architecture

Starting from the released Qwen2.5-7B layer-20 NLA, we fine-tuned its reader on table conversions of its own descriptions and adapted its reconstructor to recover activations from those tables. Constrained generation enforced three fields per row and at least six rows. To query it, we fixed the first row’s subject and relation; the model generated the value and remaining rows. It received an activation, not the source text or an answer key.

Querying an activation by fixing a subject and relation
The intended interface. Fixing two fields specifies a query; it does not establish that the generated value answers it correctly. The reconstructor reads the complete table.

Results

On 128 development examples, fixing a subject and relation already present in an earlier readout barely changed reconstruction: 60.24% → 60.12% variance explained. This checked whether forcing familiar fields preserved reconstruction, not whether the reader could answer a new question.

Reconstruction with and without forced fields
Same 128 activations, reader and reconstructor in both conditions. Variance explained is relative to the mean activation of this set.

We then assessed whether the outputs actually contained requested answers, including answers appearing outside the first row.

Query testReleased NLA, prose prefixTable-adapted NLA
192 synthetic contextual queries1/1925/192
8 factual queries from natural passagesNot tested0/8

The 192 queries covered locations, reported beliefs, and activities across 24 source templates. Manual scoring found only isolated correct answers; ambiguous cases give ranges of 1–2/192 and 3–5/192. In one credited table answer, replacing the activation with a conflicting example left the answer unchanged. These results do not establish an advantage over ordinary NLA.

The eight natural-passage queries were selected for explicit, checkable facts. All outputs had valid table structure, but none answered correctly—even elsewhere in the table. Reconstruction on this smaller set was 47.98% unforced versus 43.30% queried, using its own variance denominator. For example:

Forced subject and relationAnswer in sourceGenerated value, verbatim
InScribe Module 2 — costs£25free online course
breakfast cereal — served withmilkhistorical and scientific description of rubber
Syv Systre — countryNorwayUnited States

Preserving reconstruction did not make the reader reliably answer queries.

Discussion

The current autonomous interpretability research system still has room for improvement. For resource-intensive experiments such as the full-context NLA and queryable NLA, we still had occasional manual interventions, and our system has yet to generate a paradigm shift in interpretability research. We’re working on making this system stronger until we have boundless open-ended interpretability research able to keep up with capabilities.