Megan Kinniment

This stream will focus on the science and development of model evaluations, especially monitorability and alignment evals.

Stream overview

Here are some examples of projects I have been interested in, but I may be interested in other projects by the time this cohort starts:

  • Improve METR's general agent performance using SFT / RL / and scaffolding. (Want to rule out a big capabilities jump with a small amount of effort).
  • Models seem quite bad at judging the quality of their own solutions. Is this a capabilities bottleneck? If so, can we shortcut running our full task suite by creating a benchmark that just requires models to judge what score an agent trajectory would have gotten?
  • How can we convert model X% time horizons into 'expected speedup for person doing X task'?
  • How can we model the performance of agents on multistep tasks. (Models like e.g. toby ord's constant hazard rate model seem interesting - but don't fit with our data in some ways)
  • Building agentic scheming evals, e.g. where models can choose to sabotage products, and evade or disrupt monitoring.
  • What fundamental weaknesses are holding models back from being dangerous, and how can we test those?

Mentors

Megan Kinniment
METR
,
Member of Technical Staff
SF Bay Area
AI Control and Monitoring
Capability and Propensity Evaluations

I am a researcher at METR.

I think the development of AI is going to be a confusing time for the world. I want to help provide good evidence and methodologies for tracking AI development and risk, so humanity can make sensible decisions.

I've had different roles at different times, including leading task development and our monitoring stream. I like prototyping new kinds of evaluations. I think it's healthy to read transcripts. I'm interested in what capabilities matter for being a competent agent, and why current AI agents fall short. I feel lucky that I get to spend time building an understanding of the models.

I've previously spent time at the Centre on Long-Term Risk and FHI. Before that I studied physics at university, where I did malaria diagnostics research.

Read more

Mentorship style

I'll meet with scholars 2x/week each. I'll also be generally available async and potentially for code review.

Fellows we are looking for

Various profiles could be a good fit.

Wanted:

  • Enjoys making progress quickly, some 'productive impatience'
  • Proactivity and agency (e.g. to unblock yourself, and to ask for help when you need it)
  • Comfort doing things you haven't done before (or that nobody has done before!)
  • Basic coding skills. (E.g. git, python, bash, know what uv is)
  • Able to write reasonably clean code that other researchers can build on
  • Prior experience doing research of some kind (broadly defined, could be independent)
  • Ok with reading, or excited to read some agent transcripts
  • Interest in or curious about the models
  • Interest in understanding METR as an organization, and how we can better achieve our goals.
  • Enthusiasm is always a plus!

Can independently find collaboraters, but not required

Project selection

I'll provide a list of possible projects to pick from, and talk through the options before making a decision.

Scholars can also suggest their own projects.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Policy and Governance
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy and Governance
Policy & Governance, Strategy & Forecasting
Oxford
Theory
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Theory
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Policy and Governance
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Policy and Governance
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception