This stream focuses on empirical AI control research, including defending against AI-driven data poisoning, evaluating and attacking chain-of-thought monitorability, and related monitoring/red-teaming projects. It is well-suited to applicants already interested in AI safety with solid Python skills, and ideally prior research or familiarity with control literature/tools (e.g. Inspect/ControlArena).
Alan is Head of Autonomous Systems & Control at the UK AI Security Institute, where he works on empirical AI control and monitoring. He co-authored RepliBench, an evaluation suite measuring autonomous-replication capabilities in language-model agents.
1-hour weekly meetings for going through your research log & high level guidance. Daily updates on slack are also very useful and I typically reply within 2 days to any questions.
Essential:
You may be a good fit if you also have some of:
Not a good fit:
Collaborating with other MATS scholars.
By default I'll propose several projects for you to choose from, but you can also pitch ideas that you're interested in.
The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.