Michael Chen

Research papers (technical governance or ML) related to evaluating and mitigating dangerous AI capabilities, with a focus on what's actionable and relevant for AGI companies

Stream overview

Broad topics I am interested in include:

  • Frontier safety policies and proposing actionable improvements for companies
  • Existing frontier safety regulation, such as EU GPAI Code of Practice or SB 53
  • Dangerous capability evaluations and mitigations, especially related to loss of control
  • Making frontier safety practices more likely to be adopted in China, e.g., by analyzing relevant EU/California regulation

Mentors

Michael Chen
California Governor's Office of Emergency Services
,
AI Science Advisor / DPhil Affiliate
SF Bay Area
Policy and Governance
Capability and Propensity Evaluations
Technical AI Governance

Michael Chen works on AI policy at METR and is a part-time PhD student at Oxford in technical AI governance. Michael previously worked as a software engineer at Stripe. METR's policy team has assisted companies like Google DeepMind, Amazon, and Anthropic with developing their frontier safety policies – voluntary commitments to evaluate and mitigate severe AI risks. Besides corporate advising, Michael has provided feedback on U.S. state bills and the EU AI Act GPAI Code of Practice.

Read more

Mentorship style

I like to get daily standup messages about progress that has been made on the project, and I'm happy to provide some quick async feedback on new outputs. I'll also have weekly meetings. I'm based in Constellation in Berkeley.

Fellows we are looking for

Good writers/researchers who can work independently and autonomously! I'm looking for scholars who can ship a meaningful research output end-to-end and ideally have prior experience in writing relevant papers.

Project selection

I may assign a project, have you pick from a list of projects, or talk through project ideas with you.

Streams

The Winter 2026 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting