OpenAI Control & Oversight

We are interested in AI control and scalable oversight. I'm excited to work with scholars interested in empirical projects building and evaluating control measures and oversight techniques for LLM agents, especially those based on chain of thought monitoring. I'm also interested in the science of chain of thought monitorability, misalignment and control. An ideal project ends with a paper submitted to NeurIPS/ICML/ICLR.

Stream overview

Some example projects:

  1. Forcing AI agents to externalize their cognition, either through information bottlenecks in multi-agent architectures or training techniques
  2. Training AI agents to be more honest
  3. Training AI agents to have more monitorable chain of thought traces
  4. Agentic monitors: ones that take actions, retrieve additional context and interrogate overseen agents

Mentors

Tomek Korbak
OpenAI
,
Member of Technical Staff
SF Bay Area
Misalignment Science
AI Control and Monitoring
Alignment Training Methods

I’m a Member of Technical Staff at OpenAI working on monitoring LLM agents for misalignment. Previously, I worked on AI control and safety cases at the UK AI Security Institute and on honesty post-training at Anthropic. Before that, I did a PhD at the University of Sussex with Chris Buckley and Anil Seth focusing on RL from human feedback (RLHF) and spent time as a visiting researcher at NYU working with Ethan Perez, Sam Bowman and Kyunghyun Cho.

Read more
Micah Carroll
OpenAI
,
Member of Technical Staff, Safety Systems
SF Bay Area
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations
Structural Risk and Societal Dynamics
Alignment Training Methods

Micah is a researcher on OpenAI’s safety team interested in AI deception, scalable oversight, and monitorability. He is on leave from a UC Berkeley PhD focused on AI alignment with influenceable humans, AI manipulation from RL training, and recommender-system effects.

Read more
Miles Wang
OpenAI
,
Member of Technical Staff
SF Bay Area
Misalignment Science
AI Control and Monitoring
Biosecurity
Capability and Propensity Evaluations
Adversarial Robustness and Safeguards

Miles Wang is a researcher at OpenAI whose interests span alignment, evaluations, reasoning, and science. Wang studied computer science at Harvard before joining OpenAI in March 2024.

Read more

Mentorship style

I'll meet with mentees once a week and will be available on Slack daily. By default, I'll try to pair mentees to work on a project together.

I'll meet with mentees once a week and will be available on Slack daily.

Fellows we are looking for

An ideal mentee has a strong AI research and/or software engineering background. A mentee can be a PhD student and they can work on a paper that will be part of their thesis.

By default, I'll try to pair mentees to work on a project together.

Project selection

I'll talk through project ideas with scholar

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Strategy and Forecasting
Policy and Governance
SF Bay Area
Founding and Field-Building
London
Biosecurity
Washington, D.C.
Biosecurity
Empirical
London
Empirical
London
Empirical
SF Bay Area
Policy and Governance
SF Bay Area
Empirical
Washington, D.C.
Empirical
Washington, D.C.
Policy and Governance
Strategy and Forecasting
Policy and Governance
SF Bay Area
Policy and Governance
London
Strategy and Forecasting
Biosecurity
SF Bay Area
Founding and Field-Building
Empirical
London
Theory
Empirical
Washington, D.C.
Systems Security