Building Better Activation Oracles

MATS Fellow:

Jan Bauer, Celeste De Schamphelaere

Authors:

Jan Bauer, Celeste De Schamphelaere, Adam Karvonen, Niclas Luick, Neel Nanda

Citations

Citations

Abstract:

Activation Oracles (AOs) are promising methods for interpreting residual stream activations. However, current AOs face important issues, such as hallucinations and vagueness. Additionally, text-inversion confounds make them hard to evaluate. To this end, we improve the Activation Oracle (AO) training regime in four ways: training on on-policy rollouts, improving the conversational dataset, feeding more layers and an improvement to the injection formula. The capability improvements are marginal, but quality of life improvements are quite substantial. In addition, we open source the first comprehensive evaluation suite for AO quality, which we call AObench. Overall, we hope that our work sets a foundation that helps improve AOs and other models in the paradigm of scalable, end-to-end interpretability.

Recent research

Synthetic Persona Pretraining: Alignment from Token Zero

Authors:

Julian Minder

Date:

August 13, 2026

Citations:

Capability Provenance in Language Models: A Case Study in Social Reasoning

Authors:

Glenn Matlin, Taywon Min

Date:

August 10, 2026

Citations:

Frequently asked questions

What is the MATS Program?
How long does the program last?