MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Daniel is a professor of computer science at UIUC, where he studies the progress of AI, with a particular focus on dangerous capabilities of AI agents. His work includes:

- CVE-Bench, an award winning benchmark (SafeBench award, ICML spotlight) that is used by frontier labs and governments to measure AI agents' ability to find and exploit real-world vulnerabilities. 

- Agent Benchmark Checklist, an award winning work (Berkeley AI summit, 1st place Benchmarks & Evaluations track) that highlights major issues in existing benchmarks.

- InjecAgent, one of the first AI agent safety benchmarks, used by governments and major labs.

Focus:
实证研究
Capability and Propensity Evaluations, Biosecurity, AI Systems Security
Sara Price
Anthropic
,
Member of Technical Staff
—

Sara Price leads Alignment Training at Anthropic. Sara's research focuses on alignment training that generalizes reliably, including methods for reducing agentic misalignment.

Focus:
实证研究
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science, Alignment Training Methods

Roger Grosse is an associate professor of computer science at the University of Toronto and a member of Anthropic’s Alignment Science team, where he works on training data attribution. He earned his PhD in computer science from MIT.

Focus:
实证研究
Adversarial Robustness and Safeguards, Misalignment Science, Alignment Training Methods, Interpretability

Xander Davies is a Member of the Technical Staff at the UK AI Security Institute, where he leads the Red Teaming group, which uses adversarial ML techniques to understand, attack, and mitigate frontier AI safeguards. He is also a PhD student at the University of Oxford, supervised by Dr. Yarin Gal. He previously studied computer science at Harvard, where he founded and led the Harvard AI Safety Team.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, AI Systems Security
Maksym Andriushchenko
ELLIS Institute Tübingen
,
Principal Investigator (AI Safety and Alignment Group)
—

I am a principal investigator at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, where I lead the AI Safety and Alignment group. I also serve as chapter lead for the new edition of the International AI Safety Report chaired by Prof. Yoshua Bengio. I have worked on AI safety with leading organizations in the field (OpenAI, Anthropic, UK AI Safety Institute, Center for AI Safety, Gray Swan AI). I obtained my PhD in machine learning from EPFL in 2024 advised by Prof. Nicolas Flammarion. My PhD thesis was awarded the Patrick Denantes Memorial Prize for the best thesis in the CS department of EPFL and was supported by the Google and Open Phil AI PhD Fellowships.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Misalignment Science, Technical AI Governance, AI Systems Security
Sarah Schwettmann
Transluce
,
Co-Founder, Chief Scientist
—

I’m a Research Scientist in MIT CSAIL with the MIT-IBM Watson AI Lab. I did my PhD in Brain and Cognitive Sciences at MIT, as an NSF Fellow working with Josh Tenenbaum and Antonio Torralba. My work investigates representations underlying intelligence in artificial (and previously, biological) neural networks.

Focus:
实证研究
AI Control and Monitoring, Adversarial Robustness and Safeguards, Misalignment Science, Interpretability

Jacob Hilton is a researcher at the Alignment Research Center (ARC), a nonprofit working on the theoretical foundations of mechanistic interpretability. He previously worked at OpenAI on reinforcement learning from human feedback, scaling laws and interpretability. His background is in pure mathematics, and he holds a PhD in set theory from the University of Leeds, UK.

Focus:
Theory
Interpretability, Theoretical Alignment and Formal Methods

Eli is working on AI scenario forecasting with the AI Futures Project, where he co-authored AI 2027. He advises Sage, an organization he cofounded that works on AI Digest (interactive AI explainers) and forecasting tools. He previously worked on the AI-powered research assistant Elicit.

Focus:
战略与预测
Policy and Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics

As AI Science Advisor to the California Governor’s Office of Emergency Services (Cal OES), Michael Chen advises senior leadership on frontier AI safety and risk assessment, with a particular emphasis on critical safety incidents, AI and cyber defense, and risk from developers’ internal deployment of AI, such as sabotage by AI agents and automated AI R&D. As Science Advisor, he coordinates with AI governance leads across California’s state government and facilitates cross-sector collaboration with the academic research community, the private sector, and community and nonprofit organizations.

​

Michael previously worked on evaluations-based AI governance at METR, an independent California-based nonprofit evaluator of autonomous AI agent capabilities and risks. He advised leading AI developers on frameworks for assessing, mitigating, and transparently disclosing catastrophic AI risks. He also assisted with third-party evaluations, including a review of a developer’s report assessing sabotage risk from AI agents, and contributed to a catalog of incidents in which AI agents acted beyond their operators’ intent. Michael has engaged with U.S. government bodies on frontier AI evaluation, including the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST), and conducted research at UC Berkeley’s Center for Human-Compatible AI on learning human preferences for language models. His research and commentary on AI have been covered in outlets such as Time, The Guardian, and MIT Technology Review. Michael is a part-time PhD student at the University of Oxford and an affiliate of the Oxford Martin AI Governance Initiative.

Focus:
实证研究
Capability and Propensity Evaluations, Policy and Governance, Technical AI Governance
Tomek Korbak
OpenAI
,
Member of Technical Staff
—

I’m a Member of Technical Staff at OpenAI working on monitoring LLM agents for misalignment. Previously, I worked on AI control and safety cases at the UK AI Security Institute and on honesty post-training at Anthropic. Before that, I did a PhD at the University of Sussex with Chris Buckley and Anil Seth focusing on RL from human feedback (RLHF) and spent time as a visiting researcher at NYU working with Ethan Perez, Sam Bowman and Kyunghyun Cho.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Alignment Training Methods

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?