MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Bryce Cai
SecureBio
,
Senior Research Engineer

Bryce Cai works on bio evaluations as part of SecureBio's AI team. Cai is a co-author of ABC-Bench, a suite of tasks measuring the biosecurity-relevant capabilities of AI agents on DNA design and laboratory automation.

Focus:
Biosecurity
Biosecurity

Leonie is a Policy Director at Secure AI Project. Previously, she co-led the General-Purpose AI Code of Practice taskforce at the European AI Office, and before that had various roles at the Centre for the Governance of AI, the Institute for Law & AI, and the German government. Leonie is affiliated with the University of Berkeley, and has a background in law.

Focus:
Policy and Governance
Jacob Kaffey
SecureBio
,
Research Engineer

Jacob Kaffey is a Research Engineer at SecureBio with a background in Genetics and Computer Science, specializing in applying machine learning to healthcare. Diverse experience in data engineering, MLOps, and industry research.

Focus:
Biosecurity
Biosecurity
Lucius Bushnaq
Goodfire AI
,
Member of Technical Staff

Lucius Bushnaq is a Research Scientist at Goodfire.

He works on parameter decomposition methods for interpretability, such as Attribution-based Parameter DecompositionStochastic Parameter Decomposition, and adVersarial Parameter Decomposition. Alongside this, he works on learning theory and the theory behind interpretability, for example theoretical frameworks for computation in superposition and connections between singular learning theory and algorithmic information theory.

Previously, Lucius was a member of the interpretability team at Apollo Research, where he worked on the Local Interaction Basis and degeneracy in the loss landscape. He holds a PhD in physics.

Focus:
Empirical

Buck is the CEO of Redwood Research.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science, Forecasting and Strategy
Ethan Perez
Anthropic
,
Member of Technical Staff

Ethan Perez is a researcher at Anthropic, where he leads a team working on AI control, adversarial robustness, and other areas of AI safety research. His interests span many areas of LLM safety; he's previously led work on sleeper agentsred-teaming language models with language modelsdeveloping AI safety via debate using LLMs, and demonstrating and improving unfaithfulness in chain of thought reasoning. Read more on his website.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science
Samuel Marks
Anthropic
,
Member of Technical Staff

Sam leads the Cognitive Oversight subteam of Anthropic's Alignment Science team. Their goal is to be able to oversee AI systems not based on whether they have good input/output behavior, but based on whether there's anything suspicious about the cognitive processes underlying those behaviors. For example, one in-scope problem is "detecting when language models are lying, including in cases where it's difficult to tell based solely on input/output". His team is interested in both white-box techniques (e.g. interpretability-based techniques) and black-box techniques (e.g. finding good ways to interrogate models about their thought processes and motivations). For more flavor on this research direction, see his post here.

Focus:
Empirical
AI Control and Monitoring, Misalignment Science, Interpretability
Neel Nanda
Google DeepMind
,
Staff Research Scientist

Neel leads the mechanistic interpretability team at Google DeepMind, trying to use the internals of models to understand them better, and use this to make them safer - eg detecting deception, understanding concerning behaviours, and monitoring deployed systems for harmful behaviour.

Since mid 2024, Neel has become more pessimistic about ambitious mechanistic interpretability, and more optimistic that pragmatic approaches can add a lot of value. He's doing less work on basic science, and working more on model biology work, and work applying interpretability to real-world safety problems like monitoring.

He has spent far too much time having MATS scholars, and has about 50 alumni - he's excited to take on even more!

Focus:
Empirical
AI Control and Monitoring, Interpretability

Marius Hobbhahn is the CEO of Apollo Research, where he also leads the monitoring team. Apollo is an AI safety research organization focused on scheming, evals and control/monitoring. He is a TIME100 in AI2025 recipient. Prior to starting Apollo, Marius did a PhD in Bayesian ML and worked on AI forecasting at Epoch.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, Forecasting and Strategy
Fabien Roger
Anthropic
,
Member of Technical Staff

Fabien Roger is an AI safety researcher at Anthropic and previously worked at Redwood Research. Fabien’s research focuses on AI control and dealing with alignment faking.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, Alignment Training Methods

Frequently asked questions

What is the MATS Program?
Who are the MATS Mentors?
What are the key dates of the MATS Program?
Who is eligible to apply?
How does the application and mentor selection process work?