Shi Feng

The stream will focus on conceptual, empirical, and theoretical work on scalable oversight and control. This includes but is not limited to creating model organisms for specific failure modes, designing training procedures against them, and making progress on subproblems involved in safety cases.

Stream overview

  • Realistic model organisms of deception and collusion
  • Easy-to-hard generalization of monitors
  • Legibility and easy-to-hard generalization of scalable oversight protocols
  • Metacognition

Mentors

Shi Feng
George Washington University
,
Assistant professor
New York City, Washington, D.C.
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations

Shi Feng leads a research group working on oversight and control. He is an assistant professor at George Washington University. Prior to that, he was a postdoc in the NYU Alignment Research Group under Sam Bowman. He currently focuses on deception and collusion, with an emphasis on propensity and evaluation realism.

Read more

Mentorship style

  • Mentors can meet with scholars 1.5x/week on average.
  • Mentors will focus on unbottlenecking collaborators as quickly as possible and will check slack messages every few hours.

Fellows we are looking for

  • Interested in hybrid (conceptual, empirical, theoretical) work on scalable oversight / control methods and threat modeling against them.
  • The ability to articulate the crux of a (proposed) work and translate that into empirical experiments: what are the hidden assumptions? what hypothetical finding makes the idea more or less promising?
  • The ability to look at the data and do careful manual qualitative analysis, e.g., reading debate transcripts.
  • Comfortable with building scaffolds and various post-training processes.
  • Human evaluation experience is a plus.

Scholars will collaborate with people involved in the group but can also find new collaborators.

Project selection

A research agenda document will be shared ahead of time with a short list of project ideas. The scholars can also brainstorm and pitch ideas that are aligned with the research agenda. We will decide on assignments in week 2.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception