Michael Chen

Research papers (technical governance or ML) related to evaluating and mitigating dangerous AI capabilities, with a focus on what's actionable and relevant for AGI companies

Stream overview

Broad topics I am interested in include:

  • Frontier safety policies and proposing actionable improvements for companies
  • Existing frontier safety regulation, such as EU GPAI Code of Practice or SB 53
  • Dangerous capability evaluations and mitigations, especially related to loss of control
  • Making frontier safety practices more likely to be adopted in China, e.g., by analyzing relevant EU/California regulation
Mentorship style:

Standard (1-2 hours of weekly 1:1s)

Location during program:

SF Bay Area

London location preference:

No preference for this location

Berkeley location preference:

Strong preference

Mentors

Michael Chen
California Council on Science & Technology
,
AI Science Advisor / DPhil Affiliate
Policy and Governance
Capability and Propensity Evaluations
Technical AI Governance

As AI Science Advisor to the California Governor’s Office of Emergency Services (Cal OES), Michael Chen advises senior leadership on frontier AI safety and risk assessment, with a particular emphasis on critical safety incidents, AI and cyber defense, and risk from developers’ internal deployment of AI, such as sabotage by AI agents and automated AI R&D. As Science Advisor, he coordinates with AI governance leads across California’s state government and facilitates cross-sector collaboration with the academic research community, the private sector, and community and nonprofit organizations.

Michael previously worked on evaluations-based AI governance at METR, an independent California-based nonprofit evaluator of autonomous AI agent capabilities and risks. He advised leading AI developers on frameworks for assessing, mitigating, and transparently disclosing catastrophic AI risks. He also assisted with third-party evaluations, including a review of a developer’s report assessing sabotage risk from AI agents, and contributed to a catalog of incidents in which AI agents acted beyond their operators’ intent. Michael has engaged with U.S. government bodies on frontier AI evaluation, including the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST), and conducted research at UC Berkeley’s Center for Human-Compatible AI on learning human preferences for language models. His research and commentary on AI have been covered in outlets such as Time, The Guardian, and MIT Technology Review. Michael is a part-time PhD student at the University of Oxford and an affiliate of the Oxford Martin AI Governance Initiative.

Read more

Fellows we are looking for

Good writers/researchers who can work independently and autonomously! I'm looking for scholars who can ship a meaningful research output end-to-end and ideally have prior experience in writing relevant papers.

Project selection

I may assign a project, have you pick from a list of projects, or talk through project ideas with you.