We will continue working on black-box monitors for scheming in complex agentic settings, building on the success of the previous stream. Concretely, we will work on scaling our datasets and fine-tuning efforts, as described in the scalable monitoring agenda
Most likely the next projects will be about automated iterated red-team vs. blue-team games. We are currently training the blue team. We will then train the red-team and within this stream, we will try and close the loop to train them both synchronously.
The entire stream will be dedicated to building high-quality monitors for scheming.
In our last two streams, we first investigated constitutional monitors which resulted in an ICML conference paper
Then we investigated fine-tuned deliberative monitors which is most likely going to be accepted as a NeurIPS conference paper
For the foreseeable future, we will be following the scalable monitoring agenda
Standard (1-2 hours of weekly 1:1s)
London
Indifferent
Strong preference
Marius Hobbhahn is the CEO of Apollo Research, where he also leads the monitoring team. Apollo is an AI safety research organization focused on scheming, evals and control/monitoring. He is a TIME100 in AI2025 recipient. Prior to starting Apollo, Marius did a PhD in Bayesian ML and worked on AI forecasting at Epoch.
You will work on subprojects of black box monitoring. See here for details.