MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Sambhav is a research associate on the Frontier Security team at IAPS, where he focuses on AI deployments in defense and national security, standards for internally deployed models, and threat modeling.   

  

Before joining IAPS, Sambhav served as Co-director of the Cambridge AI Safety Hub, where he ran the MARS (Mentorship for Alignment Research Students) program.

Focus:
政策与治理
Capability and Propensity Evaluations, Policy and Governance, Technical AI Governance

Jan is a researcher at the Institute for AI Policy and Strategy (IAPS), where he works on technical AI governance to reduce catastrophic risk from AI. He is currently threat modelling how national security uses of frontier AI could go really badly and developing honeypots to secure internal AI agents.

​

Before joining IAPS, he was a GovAI winter fellow, a Pivotal fellow and a PhD candidate at the CISPA Helmholtz Center for Information Security. His past research spans Interpretability, ML security and AI Alignment. He holds a BA in Information Systems and a MSc in CS.

Focus:
政策与治理
AI Control and Monitoring, Alignment Training Methods, Technical AI Governance, Interpretability, AI Systems Security

Alex is a co-founder of [Resolution](https://resolution.org/), a research nonprofit using empirics, theory, and automation to get to higher confidence in alignment. He previously worked at the UK AI Security Institute, where he led strategy and operations for the £30m [Alignment Project](https://alignmentproject.aisi.gov.uk/). Before that, he built [80,000 Hours](https://80000hours.org/)’ automated headhunting product, which has made 50+ placements in AI safety and governance. He also co-founded [LASR](https://www.lasrlabs.org/) and [Leaf](https://leaf.courses/), and co-authored the original [Effective Altruism introductory programme curriculum](https://www.effectivealtruism.org/courses/introductory-program).

Focus:
创业与领域建设
Founding and Field-Building
Samuel Hammond
Foundation for American Innovation
,
Director of Artificial Intelligence & Chief Economist
—

Samuel Hammond is director of Artificial Intelligence Policy and chief economist at the Foundation for American Innovation, where his research focuses on artificial intelligence and the institutional impact of emerging technologies. He previously worked as the director of social policy for the Niskanen Center, where he remains a senior fellow; as an economist for the Government of Canada specializing in regional economic development; and as a graduate research fellow for the Mercatus Center at George Mason University.

​

Sam received a BA in economics from Saint Mary’s University and an MA in economics from George Mason University and Carleton University.

Focus:
政策与治理
Policy and Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics

Buck is the CEO of Redwood Research.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Forecasting and Strategy
Ethan Perez
Anthropic
,
Member of Technical Staff
—

Ethan Perez is a researcher at Anthropic, where he leads a team working on AI control, adversarial robustness, and other areas of AI safety research. His interests span many areas of LLM safety; he's previously led work on sleeper agents, red-teaming language models with language models, developing AI safety via debate using LLMs, and demonstrating and improving unfaithfulness in chain of thought reasoning. Read more on his website.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science
Samuel Marks
Anthropic
,
Member of Technical Staff
—

Sam leads the Cognitive Oversight subteam of Anthropic's Alignment Science team. Their goal is to be able to oversee AI systems not based on whether they have good input/output behavior, but based on whether there's anything suspicious about the cognitive processes underlying those behaviors. For example, one in-scope problem is "detecting when language models are lying, including in cases where it's difficult to tell based solely on input/output". His team is interested in both white-box techniques (e.g. interpretability-based techniques) and black-box techniques (e.g. finding good ways to interrogate models about their thought processes and motivations). For more flavor on this research direction, see his post here.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Interpretability
Neel Nanda
Google DeepMind
,
Staff Research Scientist
—

Neel leads the mechanistic interpretability team at Google DeepMind, trying to use the internals of models to understand them better, and use this to make them safer - eg detecting deception, understanding concerning behaviours, and monitoring deployed systems for harmful behaviour.

​

Since mid 2024, Neel has become more pessimistic about ambitious mechanistic interpretability, and more optimistic that pragmatic approaches can add a lot of value. He's doing less work on basic science, and working more on model biology work, and work applying interpretability to real-world safety problems like monitoring.

​

He has spent far too much time having MATS scholars, and has about 50 alumni - he's excited to take on even more!

Focus:
实证研究
AI Control and Monitoring, Interpretability

Marius Hobbhahn is the CEO of Apollo Research, where he also leads the monitoring team. Apollo is an AI safety research organization focused on scheming, evals and control/monitoring. He is a TIME100 in AI2025 recipient. Prior to starting Apollo, Marius did a PhD in Bayesian ML and worked on AI forecasting at Epoch.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, Forecasting and Strategy
Fabien Roger
Anthropic
,
Member of Technical Staff
—

Fabien Roger is an AI safety researcher at Anthropic and previously worked at Redwood Research. Fabien’s research focuses on AI control and dealing with alignment faking.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, Alignment Training Methods

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?