Stephen Casper (Cas)

I (Cas) work on a range of projects from technical safeguards to technical governance. This stream follows an academic collaboration model and work will likely focus on technical topics in AI governance.

Stream overview

Technical work: Making safeguards 'run deep', including safeguards and risk management for open-weight models.

Governance work: Critical review of industry self-governance, critical review of national AI governance institutes, open-weight model governance, predicting and mitigating future AI incidents.

Mentors

Stephen Casper (Cas)
Harvard
,
Assistant Professor
Boston
Policy and Governance
Technical AI Governance
Alignment Training Methods
Adversarial Robustness and Safeguards

Stephen "Cas" Casper is a computer scientist and an Assistant Professor of Public Policy at the Harvard Kennedy School and a Faculty Affiliate of the Harvard School of Engineering and Applied Sciences. Prior to joining Harvard, he completed his PhD at MIT and did a research residency with the UK AI Security Institute. He is a writer for the International AI Safety Report and a lead writer for the Singapore Consensus. His research has been recognized with a Hoopes Prize, an ML Safety Workshop best paper award, a BioSafeGenAI best paper runner-up, a GenLaw spotlight paper award, a TMLR outstanding paper finalist distinction, and a handful of mentions in news articles and newsletters. Find him on Google ScholarTwitter (sorry), BlueSky, and LinkedIn.

Read more

Mentorship style

2-3 meetings per week plus regular messaging and collaborative writing.

Fellows we are looking for

Green flags include:

  • Research experience and taste
  • Demonstrated initiative in ideating and leading prior projects
  • Demonstrated ability to 'succeed even when not set up to succeed' in past research/projects.
  • Taking a critical mindset about projects and impact
  • Interest in academia

This stream will follow an academic collaboration model. Scholars will be free to discuss and collaborate externally. However, scholars should also expect to work in collaboration with others in the stream.

Project selection

Mentor(s) will talk through project ideas with scholar.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception