Sarah Schwettmann, Jacob Steinhardt

We build scalable technology for AI understanding and oversight.

Stream overview

We’re building scalable, AI-backed systems for analyzing, testing, and interpreting AI agents, and using these to study behaviors like sycophancy, self-harm, and reward hacking. We’re looking for scholars who want to help us push forward this work.

Some concrete projects include: scalable, end-to-end tools for interpretability and behavior elicitation; creating robust LLM judges for Docent; scalable search and retrieval for large agent transcripts.

Mentors

Sarah Schwettmann
Transluce
,
Co-Founder, Chief Scientist
SF Bay Area
Interpretability
Misalignment Science
AI Control and Monitoring
Adversarial Robustness and Safeguards

I’m a Research Scientist in MIT CSAIL with the MIT-IBM Watson AI Lab. I did my PhD in Brain and Cognitive Sciences at MIT, as an NSF Fellow working with Josh Tenenbaum and Antonio Torralba. My work investigates representations underlying intelligence in artificial (and previously, biological) neural networks.

Read more
Jacob Steinhardt
Transluce
,
Co-founder, CEO
SF Bay Area
Interpretability
Misalignment Science
AI Control and Monitoring
Forecasting and Strategy
Capability and Propensity Evaluations

I am an Assistant Professor of Statistics and EECS at UC Berkeley, where I’m also part of BAIR and CLIMB. I am also Founder & CEO of Transluce, a non-profit research lab building open, scalable technology for understanding frontier AI systems.

Read more

Mentorship style

You will work closely with a mentor through recurring meetings (group and individual) and Slack.

Fellows we are looking for

We're looking for strong, experienced software engineers or talented researchers who can hit the ground running and iterate quickly.

ML experience is a bonus but not required.

Probably will work with collaborators from stream

Project selection

We will talk through project ideas with scholar

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception