This stream focuses on empirical AI control research, including defending against AI-driven data poisoning, evaluating and attacking chain-of-thought monitorability, and related monitoring/red-teaming projects. It is well-suited to applicants already interested in AI safety with solid Python skills, and ideally prior research or familiarity with control literature/tools (e.g. Inspect/ControlArena).
London
Alan is Head of Autonomous Systems & Control at the UK AI Security Institute, where he works on empirical AI control and monitoring. He co-authored RepliBench, an evaluation suite measuring autonomous-replication capabilities in language-model agents.
Essential:
You may be a good fit if you also have some of:
Not a good fit:
By default I'll propose several projects for you to choose from, but you can also pitch ideas that you're interested in.