MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Nicholas is a research scientist at Anthropic (previously Google DeepMind) researching adversarial machine learning; he likes to break things.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, AI Systems Security
Sam Bowman
Anthropic
,
Member of Technical Staff

Sam Bowman leads a research group working on AI alignment and welfare at Anthropic, with a particular focus on evaluation. Sam is also on leave from NYU as an Associate Prof. of Computer Science and Data Science. He has been studying neural network language models since 2012.

Focus:
Empirical
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science, AI Welfare
Joe Benton
Anthropic
,
Member of Technical Staff

Joe is a member of the Alignment Science team at Anthropic. He's currently working on scalable oversight and also has interests in control, chain-of-thought monitoring, and alignment evaluations. For some examples of recent projects, including MATS collaborations, see: https://joejbenton.com/research/.

Focus:
Empirical
AI Control and Monitoring, Adversarial Robustness and Safeguards, Misalignment Science
Yoshua Bengio
LawZero
,
Co-President and Scientific Director (LawZero) / Full Professor (UdeM) / Founder and Scientific Advisor (Mila)

Yoshua Bengio is Full Professor of Computer Science at Université de Montreal, Co-President and Scientific Director of LawZero, as well as the Founder and Scientific Advisor of Mila. He also holds a Canada CIFAR AI Chair. Considered one of the world’s leaders in Artificial Intelligence and Deep Learning, he is the recipient of the 2018 A.M. Turing Award, considered to be the "Nobel Prize of computing." He is the most cited computer scientist worldwide, and the most-cited living scientist across all fields (by total citations).

Professor Bengio is a Fellow of both the Royal Society of London and Canada, an Officer of the Order of Canada, a Knight of the Legion of Honor of France, a member of the UN’s Scientific Advisory Board for Independent Advice on Breakthroughs in Science and Technology, and chairs the International AI Safety Report.

Focus:
Empirical
AI Control and Monitoring, Policy and Governance, Agent Foundations, Theoretical Alignment and Formal Methods
Alex Turner
Independent
,
Independent researcher

Alex is currently working on training invariants into model behavior. In the past, he formulated and proved the power-seeking theorems, co-formulated the shard theory of human value formation, and proposed the Attainable Utility Preservation approach to penalizing negative side effects.

Highlighted outputs from past streams:

Focus:
Empirical
Alignment Training Methods, Interpretability, Agent Foundations

Daniel is working on forecasting detailed AI scenarios with Eli Lifland, Thomas Larsen, Jonas Vollmer, and Romeo Dean.

Focus:
Strategy and Forecasting
Policy and Governance, Technical AI Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics

Thomas is a researcher at the AI Futures Project. He was a co-author on the widely read AI 2027 scenario forecast. He previously founded the Center for AI Policy, an AI safety advocacy organization, and worked on AI safety research at the Machine Intelligence Research Institute.

Focus:
Strategy and Forecasting
Policy and Governance, Forecasting and Strategy, Structural Risk and Societal Dynamics
Trenton Bricken
Anthropic
,
Member of Technical Staff

I'm a Member of Technical Staff on the Alignment Science team at Anthropic. I'm currently enabling Claude to automatically audit and detect misalignment.

About me

  • I have a PhD in Systems Biology from Harvard. My thesis was on "Sparse Representations in Biological and Artificial Neural Networks" in the Kreiman Lab with support from the NSF Graduate Research Fellowship. I also spent time at the Berkeley Redwood Center for Theoretical Neuroscience as a visiting researcher.
  • I graduated from Duke University in May 2020 with a self-made major in "Minds and Machines: Biological and Artificial Intelligence". I was lucky to attend as a Robertson Scholar, which provided full funding during all four years, including summer experiences.
  • At Duke, I spent a year doing research in Dr. Michael Lynch's Lab attempting to use machine learning to design new CRISPR guide RNAs for safer, more effective genome editing. Afterwards, I was affiliated with Dr. Debora Marks's Lab at Harvard Medical School applying deep learning to protein design. I also contributed to the IARPA Fun GCAT and DARPA Biostasis programs.
Focus:
Empirical
AI Control and Monitoring, Misalignment Science, Interpretability
Tyler Tracy
Redwood Research
,
Member of Technical Staff

Tyler is an AI Safety Researcher and member of technical staff at Redwood Research.

Focus:
Empirical
AI Control and Monitoring

Daniel is a professor of computer science at UIUC, where he studies the progress of AI, with a particular focus on dangerous capabilities of AI agents. His work includes:

  • CVE-Bench, an award winning benchmark (SafeBench award, ICML spotlight) that is used by frontier labs and governments to measure AI agents' ability to find and exploit real-world vulnerabilities.
  • Agent Benchmark Checklist, an award winning work (Berkeley AI summit, 1st place Benchmarks & Evaluations track) that highlights major issues in existing benchmarks.
  • InjecAgent, one of the first AI agent safety benchmarks, used by governments and major labs.

Focus:
Empirical
Capability and Propensity Evaluations, Adversarial Robustness and Safeguards, Biosecurity, AI Systems Security

Frequently asked questions

What is the MATS Program?
Who are the MATS Mentors?
What are the key dates of the MATS Program?
Who is eligible to apply?
How does the application and mentor selection process work?