MATS mentors are advancing the frontiers of AI alignment, transparency, and security

Stephen "Cas" Casper is a computer scientist and an Assistant Professor of Public Policy at the Harvard Kennedy School and a Faculty Affiliate of the Harvard School of Engineering and Applied Sciences. Prior to joining Harvard, he completed his PhD at MIT and did a research residency with the UK AI Security Institute. He is a writer for the International AI Safety Report and a lead writer for the Singapore Consensus. His research has been recognized with a Hoopes Prize, an ML Safety Workshop best paper award, a BioSafeGenAI best paper runner-up, a GenLaw spotlight paper award, a TMLR outstanding paper finalist distinction, and a handful of mentions in news articles and newsletters. Find him on Google Scholar, Twitter (sorry), BlueSky, and LinkedIn.

Focus:
政策与治理
Adversarial Robustness and Safeguards, Alignment Training Methods, Policy and Governance, Technical AI Governance

I am a research scientist on the AGI Safety & Alignment team at Google DeepMind. I focus on deceptive alignment and AI control, particularly [scheming propensity evaluations](https://arxiv.org/abs/2605.29729). My past research includes dangerous capability evals, power-seeking incentives, specification gaming, and avoiding side effects.

Focus:
实证研究
AI Control and Monitoring, Capability and Propensity Evaluations, Misalignment Science
Mauricio Baker
RAND, University of Oxford
,
Technical AI Policy Research Scientist; DPhil (PhD) student
—

Mauricio researches AI policy at RAND and Oxford. His work has focused on verification of international agreements on AI. He’s more broadly interested in technical AI governance. Previously, Mauricio contracted with OpenAI and did a master's in Computer Science at Stanford University.

Focus:
系统安全
Policy and Governance, Technical AI Governance, AI Systems Security
Alex Mallen
Redwood Research
,
Member of Technical Staff
—

Alex is a member of technical staff at Redwood Research.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Forecasting and Strategy, Interpretability, Theoretical Alignment and Formal Methods
Abram Demski
AFFINE
,
Research Scientist
—

Abram Demski is an AI Safety researcher specializing in Agent Foundations, best known for Embedded Agency (co-written with Scott Garrabrant). His overall approach primarily involves deconfusion research in relation to various concepts related to AI risks, including agency, optimization, trust, meaning, understanding, interpretability, and computational uncertainty (more commonly but less precisely known as bounded rationality). More specifically, his recent work focuses on modeling trust, with the objective of clarifying conditions under which humans can justifiably trust AI.

Focus:
Theory
Agent Foundations, Theoretical Alignment and Formal Methods
James Lucassen
Redwood Research
,
Member of Technical Staff
—

James is a member of technical staff at Redwood Research.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science
Aryan Bhatt
Redwood Research
,
Member of Technical Staff
—

Aryan is a senior member of technical staff at Redwood Research.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science
Jack Lindsey
Anthropic
,
Member of Technical Staff
—

Hi, I'm Jack! I'm interested in understanding the cognition of modern language models, so that we can make them more reliable and aligned with human values. Currently, I lead the "Model Psych" team at Anthropic. We study the internal basis of higher-level cognitive phenomena in LLMs, like introspection, situational awareness, personas, and representations of emotion. We apply these techniques to audit Anthropic’s production models, for instance by monitoring their neural activity for signatures of deception, manipulation, or awareness of being evaluated. Previously, I did my PhD in the Center for Theoretical Neuroscience at Columbia University. For a list of my publications, see my Google Scholar profile.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science, Interpretability
Vivek Hebbar
Redwood Research
,
Member if Technical Staff
—

Vivek is a member of technical staff at Redwood Research.

Focus:
实证研究
AI Control and Monitoring, Misalignment Science

I previously worked on the alignment team at DeepMind, and on the governance team at OpenAI. I'm currently an independent researcher focusing on multi-agent intelligence. My research is in the tradition of natural philosophy; I'm trying to develop vague intuitive concepts (like trust, identity, and integrity) to the point where they can serve as seeds for new scientific paradigms.

Focus:
Theory
Forecasting and Strategy, Structural Risk and Societal Dynamics, Agent Foundations

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?