Alan Cooney

This stream focuses on empirical AI control research, including defending against AI-driven data poisoning, evaluating and attacking chain-of-thought monitorability, and related monitoring/red-teaming projects. It is well-suited to applicants already interested in AI safety with solid Python skills, and ideally prior research or familiarity with control literature/tools (e.g. Inspect/ControlArena).

Stream overview

Mentors

Alan Cooney
UK AISI
,
Researcher
London
AI Control and Monitoring
Capability and Propensity Evaluations

Alan is Head of Autonomous Systems & Control at the UK AI Security Institute, where he works on empirical AI control and monitoring. He co-authored RepliBench, an evaluation suite measuring autonomous-replication capabilities in language-model agents.

Read more

Mentorship style

1-hour weekly meetings for going through your research log & high level guidance. Daily updates on slack are also very useful and I typically reply within 2 days to any questions.

Fellows we are looking for

Essential:

  • Existing interest in AI safety
  • Programming experience with Python/similar

You may be a good fit if you also have some of:

  • Research experience: prior AI research experience (on any topic)
  • Strong Python skills (essential for CoT monitorability eval projects)
  • Familiarity with the Control literature - see here
  • Familiarity with Inspect/ControlArena

Not a good fit:

  • Scholars primarily interested in conceptual rather than empirical research

Collaborating with other MATS scholars.

Project selection

By default I'll propose several projects for you to choose from, but you can also pitch ideas that you're interested in.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception