Theory

The hardest problems in AI safety may not be solvable with experiments alone. Instead, they may require the kind of foundational thinking in mathematics and philosophy that gives the field something solid to build on. Streams in this track work on agent foundations, formal models of trust and agency, mechanistic interpretability theory, and AI welfare. We're looking for researchers with deep mathematical maturity who want to tackle the problems that will still matter when AI systems are far more capable than they are today.

Application process

  • Stage 1: Complete the general application
  • Stage 2: Stream selection questions alongside possibly reference requests
  • Stage 3: Interviews and work-tests

Theory track overview

This track works on problems where the goal is durable conceptual progress rather than experimental results on today's models. The assumption is that some of the hardest alignment issues, such as questions around agency, optimization, trust, and the structure of cognition, will not be settled by simply scaling current empirical techniques, and that mathematical and philosophical foundations will matter significantly when AI systems are much more capable than they are now. Projects here cover agent foundations, formal models of trust and agency, mechanistic interpretability theory, and AI welfare. Methods rely mainly on reasoning and proof, though some work intersects with empirical interpretability or formal verification.

We are looking for fellows with serious mathematical maturity and a willingness to sit with problems where the right formalization is itself part of the work. Essential traits are research independence (theory questions are open-ended and require self-direction), fluency with formal reasoning (proofs, probability, type theory, dynamical systems, or analogous), and the ability to write clearly about abstract ideas. Strong candidates have come from mathematics, theoretical computer science, theoretical physics, formal philosophy, and economic theory, but a fellow’s background is not as significant as their demonstrated ability to do hard formal work, which can ideally be documented in written output we can read.

Fellows are matched to mentors based on fit and produce concrete artifacts (e.g., papers, technical reports, conceptual write-ups, or formal results) by the end of the program . Target audiences for the work produced in this track include the agent foundations and alignment theory communities, alignment-relevant teams at frontier labs, and academic venues for formal work. Theory outputs typically have longer time horizons than empirical ones, and we expect many fellows to continue refining results past the conclusion of the program.

Theory track streams

Agent Foundations research focused on clarifying conditions under which humans can justifiably trust artificial intelligence systems. When should one boundedly rational learning-theoretic process come to trust another?

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

The Alignment Research Center is a small non-profit research group based in Berkeley, California, that is working on a systematic and theoretically grounded approach to mechanistically explaining neural network behavior. We are interested in fellows with a strong math background and mathematical maturity. If you'd be excited to work on the research direction described in this blog post – then we'd encourage you to apply!

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

Theory of change: Soon, most important work will be done by AI. AI is going to increasingly advise people and help with important things, many of which are time-sensitive and path dependent, e.g., work on alignment/safety (including various things like how LLMs should behave given that they’re very persuasive); how to think about acausal trade; how to organize society. It seems good for AI to do well at those things.

Of course, a lot of the relevant skills for doing well at these tasks are the same skills that cause AI risk and that AI companies work on (and are incentivized to work on) by default; like coding, some kinds of forecasting, etc.

We want to make models better at things that are net positive for the future, but that likely won’t benefit much from said default training (or perhaps will even be made worse by such training – e.g., via sycophancy).

In practice, a lot of the tasks that we’re interested in from this perspective are what we call “conceptual”: tasks that are hard to verify and don't have clear ground truth but where we nonetheless feel like we can make progress through argument and reason.

You can visit conceptualreasoning.ai to get a sense of our work to date.

We also take a keen interest in projects directly aimed at making future acausal interactions go well.

Read more
Desired fellow characteristics

Projects in this stream will be on AI welfare and moral status; more specifically, on what it takes to be a moral patient and how we can determine whether AI systems meet the conditions. I'm looking for applicants who have ideas about these topics and are motivated to explore them in more detail.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

Existing frameworks for understanding intelligent agency don't do a great job at describing multi-agent dynamics (e.g. agents recursively modeling each other, merging with each other, threatening each other, etc). Most work in my stream aims (implicitly or explicitly) to move towards a multi-agent understanding of intelligence.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

Frequently asked questions

What is the MATS Program?
Who are the MATS Mentors?
What are the key dates of the MATS Program?
Who is eligible to apply?
How does the application and mentor selection process work?