Richard Ngo

My MATS fellows will do philosophical thinking about multi-agent intelligence and how agents change their values. This will likely involve trying to explore and synthesize ideas from game theory, signaling theory, reinforcement learning, and other related domains.

Stream overview

I'd like theoretically-inclined scholars to work with me towards building up a theory of coalitional agency. This work will be highly abstract, philosophical and exploratory; it's mainly suited for people with strong mathematical backgrounds who've read some of my existing writings on this topic and find them interesting.

Mentors

Richard Ngo
Independent
,
Researcher
SF Bay Area
Agent Foundations
Forecasting and Strategy
Structural Risk and Societal Dynamics

I previously worked on the alignment team at DeepMind, and on the governance team at OpenAI. I'm currently an independent researcher focusing on multi-agent intelligence. My research is in the tradition of natural philosophy; I'm trying to develop vague intuitive concepts (like trust, identity, and integrity) to the point where they can serve as seeds for new scientific paradigms.

Read more

Mentorship style

I'll come meet scholars in person around 2 days a week on average. On those days I'll be broadly available for discussions and brainstorming. On other days scholars can message me for guidance (though I'd prefer to spend most of my effort on this during the in-person days).

Fellows we are looking for

My main criterion for selecting scholars will be clarity of reasoning.

You will probably will work with other scholars in the stream.

Project selection

I will talk through project ideas with the scholar.

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

Grand Rapids
Theory
Agent Foundations
Washington, D.C.
Systems Security
Policy and Governance
Compute Infrastructure, Policy & Governance, Security
London
Theory
Policy and Governance
Dangerous Capability Evals, Compute Infrastructure, Policy & Governance, Strategy & Forecasting
Washington, D.C.
Systems Security
Compute Infrastructure, Security
London
Control, Monitoring
London
Control, Scheming & Deception, Dangerous Capability Evals, Model Organisms, Monitoring
SF Bay Area
Security, Dangerous Capability Evals
SF Bay Area
Dangerous Capability Evals
Boston
Policy and Governance
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy and Governance
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Theory
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations