Streams in this track include hands-on research using machine- learning experiments to understand and improve model safety including AI control, interpretability, scalable oversight, evaluations, red-teaming, and robustness. This track is defined by its methods rather than any single research agenda. If your primary tool is ML engineering, this is the track for you.
The track is defined by its methodology more than by any single research agenda. Fellows run ML experiments to understand and improve the safety properties of frontier models, with work spanning interpretability, AI control, scalable oversight, evaluations, red-teaming, robustness, and model organisms of misalignment. The unifying thread is that progress comes from hands-on work with real models (training, probing, fine-tuning, measuring, etc.) rather than reasoning from first principles alone. This is the largest track in the program and the most common entry point into technical AI safety research.
We are looking for fellows whose primary tool is ML engineering, broadly construed. The essential requirement is the ability to design and run experiments on language models or other deep learning systems and iterate quickly on the results. In practice, that usually means having a solid understanding of Python (with and without AI coding tools), being comfortable with the infrastructure around running models at moderate scale, and knowing which experiments are worth running. Mission alignment is highly important, and fellows should be able to say why a given line of empirical work meaningfully reduces frontier risk, not just whether it yields a successful publication. Educational background and seniority are weighted lightly here relative to other tracks. Past cohorts have included strong fellows ranging from undergraduates to senior industry researchers.
Fellows are matched to mentors based on fit, and projects are scoped to produce concrete artifacts (i.e., papers, evaluation suites, open-source tooling, or technical reports) by the end of the program. The target audiences for the work produced in this track would include safety and alignment teams at frontier labs, governments and other evaluation organizations, and the broader ML research community.
If you are excited by this kind of work, we encourage you to apply.
This stream focuses on secret loyalties, where an LLM covertly tries to advance a principal's interests. Secret loyalties have been established as a pressing threat [1], and model organisms of narrow secret loyalties have been constructed and audited [2]. This stream aims to advance the empirical foundations of our understanding of secret loyalties. The goal being that humanity is well-equipped to deal with secret loyalty installation attempts as and when catastrophic secret loyalties become possible in the future.
We think being well-equipped looks like having sufficient security measures in place in frontier AI companies, understanding the dynamics and behaviours of secretly loyal AI systems, having effective auditing and verification protocols for secret loyalties and attempts to install them, and these protocols actually being followed by relevant stakeholders.
This coalition of mentors make up the “Anthropic Stream”. This stream spans a range of empirical research areas in AI safety on LLMs, including AI control, scalable oversight, model organisms, model internals, model welfare, security, and more. You’ll be pitched, and have the option to pitch, a variety of safety research projects, and then be matched to projects and mentors based on your interests/preferences on research and what you’d like to get out of MATS. Fellows in this stream frequently receive funding and continued mentorship after MATS to complete their research project, usually leading to a (co-)first author paper. People in this stream often end up in long-term homes for safety research after MATS (e.g. Anthropic, Redwood Research, OpenAI).
Anthropic mentors share an application, tend to collaborate and co-mentor projects together, and generally share infrastructure to streamline the fellow experience. By applying to this stream, you are being considered for all of the Anthropic mentors.
During the program, scholars meet weekly with their project mentors and collaborators. Some projects meet more often without mentors (e.g., daily standups with the peers on the project). Each project will have a primary mentor, who is also the main decision-maker on key milestones for the project and who is the default person to go to for feedback, advice, etc. Co-mentors also attend project meetings as needed and provide feedback throughout the program. Some project co-mentors can be as involved as the primary mentor.
Mentorship starts with the “Project Pitch Session” Anthropic runs at the start of the program. Fellows get ~1 week to derisk and trial projects before submitting their preferences. Starting on week 2, scholars are assigned projects where the primary mentor is whoever pitched it. Some projects are assigned co-mentors who are other supervisors who want to join the project.
We will continue working on black-box monitors for scheming in complex agentic settings, building on the success of the previous stream. Concretely, we will work on scaling our datasets and fine-tuning efforts, as described in the scalable monitoring agenda
Most likely the next projects will be about automated iterated red-team vs. blue-team games. We are currently training the blue team. We will then train the red-team and within this stream, we will try and close the loop to train them both synchronously.
We have two weekly 60-minute calls by default. Since everyone will work on the same project, these calls will be with all participants of the stream. I respond on slack on a daily basis for asynchronous messages. Scholars will have a lot of freedom for day-to-day decisions and direction setting. In the best case, you will understand the project better than me after a few weeks and have a clear vision for where it should be heading. I recommend scholars focus 100% of their work time on the project and not pursue anything on the side. I think this way people will learn the most in MATS.
You will work on subprojects of black box monitoring. See here for details.
Theory of change: Soon, most important work will be done by AI. AI is going to increasingly advise people and help with important things, many of which are time-sensitive and path dependent, e.g., work on alignment/safety (including various things like how LLMs should behave given that they’re very persuasive); how to think about acausal trade; how to organize society. It seems good for AI to do well at those things.
Of course, a lot of the relevant skills for doing well at these tasks are the same skills that cause AI risk and that AI companies work on (and are incentivized to work on) by default; like coding, some kinds of forecasting, etc.
We want to make models better at things that are net positive for the future, but that likely won’t benefit much from said default training (or perhaps will even be made worse by such training – e.g., via sycophancy).
In practice, a lot of the tasks that we’re interested in from this perspective are what we call “conceptual”: tasks that are hard to verify and don't have clear ground truth but where we nonetheless feel like we can make progress through argument and reason.
You can visit conceptualreasoning.ai to get a sense of our work to date.
We also take a keen interest in projects directly aimed at making future acausal interactions go well.
None
This stream will focus on model motivations and character, open-ended environments, and new forms of misalignment.
This stream will focus on monitoring, stress-testing safety methods, and evals, with a focus on risks from scheming AIs. Examples include (black-box) AI control techniques, white-box monitors (probes etc.), chain-of-thought monitoring/faithfulness, building evaluation environments, and stress-testing mitigations.
For each project, we will have a weekly meeting to discuss the overall project direction and prioritize next steps for the upcoming week. On a day-to-day basis, you will discuss experiments and write code with other mentees on the project (though I'm available on Slack for quick feedback between meetings or to address things that are blocking you).
I structure the program around collaborative, team-based research projects. You will work in a small team, on a project from a predefined list. I organize the 12-week program into fast-paced research sprints designed to create and keep research velocity, so you should expect regular deadlines and milestones. I will provide a more detailed schedule and set of milestones at the beginning of the program.
I am looking for scholars with strong machine learning engineering skills, as well as a background in technical research. While I’ll provide weekly guidance on research, I expect scholars to be able to run experiments and decide on low-level details fairly independently most of the time. I’ll propose concrete projects to choose from, so you should not expect to work on your own research idea during MATS. I strongly encourage collaboration within the stream, so you should expect to work in teams of 2-3 scholars on a project, hence good communication and team skills are important.
We will most likely have a joint project selection phase, where we present a list of projects (with the option for scholars to iterate on them). Afterward, each project will have at least one main mentor, but we might also co-mentor some projects.
At Fourth Eon Biosecurity we're building adaptive, AI-native safeguards across the bioengineering stack, with a focus on function-based DNA synthesis screening. Fellows in this stream will work on technical research projects at the intersection of AI safety and biosecurity, aimed at reinforcing screening and generalizing detection beyond known threat signatures. Projects span mechanistic interpretability of bio foundation models, model evaluations for biosecurity-relevant capabilities, and agentic sequence analysis workflows.
I typically schedule a standing weekly 1:1 meeting with each fellow, and also hold a weekly research group meeting. Beyond that I am available on Slack and can find additional time for calls outside of scheduled meetings.
Note that as part of our Safe and Responsible Research Framework we require fellows to sign a fellowship agreement covering confidentiality and pre-publication review for dual-use risks. This is common practice in biosecurity research and allows us to work freely together on sensitive material.
Fellows who are interested in our research area should think of potential project ideas that leverage their strengths and interests. I will work with individual fellows to identify a specific project that matches their background and interests and is aligned with our overall research direction, and to refine the scope and objectives of the project.
I'm interested in better understanding and controlling how post-training causes alignment-relevant behavior. This is a pretty broad area, and I’m open to many approaches to these problems! Potential areas of study / methods of attack might include model organisms, training run science/ablations, root causing strange behaviors, or studying how best to robustly induce behaviors or values or beliefs into models.
The MATS Program is a 10-week research fellowship designed to train and support emerging researchers working on AI alignment, transparency and security. Fellows collaborate with world-class mentors, receive dedicated research management support, and join a vibrant community in Berkeley focused on advancing safe and reliable AI. The program provides the structure, resources, and mentorship needed to produce impactful research and launch long-term careers in AI safety.
MATS mentors are leading researchers from a broad range of AI safety, alignment, governance, field-building and security domains. They include academics, industry researchers, and independent experts who guide scholars through research projects, provide feedback, and help shape each scholar’s growth as a researcher. The mentors represent expertise in areas such as:
Key dates
Application:
The main program will then run from September 28th to December 4th, with the extension phase for accepted fellows beginning in December.
MATS accepts applicants from diverse academic and professional backgrounds - from machine learning, mathematics, and computer science to policy, economics, physics, cognitive science, biology, and public health, as well as founders, operators, and field-builders without traditional research backgrounds. The primary requirements are strong motivation to contribute to AI safety and evidence of technical aptitude, research potential, or relevant operational experience. Prior AI safety experience is helpful but not required.
Applicants submit a general application, applying to various tracks (Empirical, Theory, Strategy & Forecasting, Policy & Governance, Systems Security, Biosecurity, Founding & Field-Building.
In stage 2, applicants apply to streams within those tracks as well as completing track specific evaluations.
After a centralized review period, applicants who are advanced will then undergo additional evaluations depending on the preferences of the streams they've applied to before doing final interviews and receiving offers.
For more information on how to get into MATS, please look at this page.