Streams in this track include hands-on research using machine learning experiments to understand and improve model safety including AI control, interpretability, scalable oversight, evaluations, red-teaming, and robustness. This is the largest track in the program and is defined by its methods rather than any single research agenda. If your primary tool is ML engineering, this is your track.
The track is defined by its methodology more than by any single research agenda. Fellows run ML experiments to understand and improve the safety properties of frontier models, with work spanning interpretability, AI control, scalable oversight, evaluations, red-teaming, robustness, and model organisms of misalignment. The unifying thread is that progress comes from getting hands on real models (training, probing, fine-tuning, measuring) rather than reasoning from first principles alone. This is the largest track in the program and the most common entry point into technical AI safety research.
We are looking for fellows whose primary tool is ML engineering, broadly construed. The essential requirement is the ability to design and run experiments on language models or other deep learning systems and iterate quickly on the results. In practice that usually means strong Python (with and without AI coding tools), comfort with the infrastructure around running models at moderate scale, and enough research taste to know which experiments are worth running. Mission alignment matters: fellows should be able to say why a given line of empirical work meaningfully reduces frontier risk, not just whether it yields a successful publication. Educational background and seniority are weighted lightly here relative to other tracks. Past cohorts have included strong fellows ranging from undergraduates to senior industry researchers.
Fellows are matched to mentors based on fit, and projects are scoped to produce concrete artifacts by program end: papers, evaluation suites, open-source tooling, or technical reports. Target audiences include safety and alignment teams at frontier labs, governments and other evaluation organizations, the broader ML research community.
This stream will focus on monitoring, stress-testing safety methods, and evals, with a focus on risks from scheming AIs. Examples include (black-box) AI control techniques, white-box monitors (probes etc.), chain-of-thought monitoring/faithfulness, building evaluation environments, and stress-testing mitigations.
This varies a bit by mentor, but generally we expect fellows to autonomously drive their project forward (alongside their teammate(s)) and meet with their mentor once a week.
Fellows are responsible for:
Mentors are responsible for:
In addition to weekly meetings, you may set up ad hoc meetings with your mentor, reach them on Slack, or send them proposals / paper drafts for review. All meetings will be at times compatible with the BST / CEST timezones.
We're looking for fellows who want to dedicate their careers to making transformative AI go well, and who have strong technical skills. You don't need to be exceptional at every technical skill below, but you should have basic competence across the list and be excellent (or improving fast) in some of them.
Mission orientation
Technical skills
Working style
Team player: You impartially consider others' ideas, disagree respectfully, admit when you're wrong, and commit to the team's direction once decided. You are happy to work in teams of 2-3.
Close to program start, we will share a list of proposed projects, each tied to a primary mentor (some projects may also have a secondary mentor). You will then have some time to think, look up relevant literature, ask mentors questions and talk to other fellows. Then we will send out a form for you to indicate your top N projects, ranked, as well as any teammate or mentor preferences. We will then optimise fellow allocation into teams and projects to maximally satisfy everyone’s preferences.
We are excited to supervise projects:
Essential knowledge:
Essential experience:
Desired experience:
Bonus:
Lee's stream will focus primarily on improving mechanistic interpretability methods for reverse-engineering neural networks.
Mentorship looks like a 1 h weekly meeting by default with approximately daily slack messages in between. Usually these meetings are just for updates about how the project is going, where I’ll provide some input and steering if necessary and desired. If there are urgent bottlenecks I’m more than happy to meet in between the weekly interval or respond on slack in (almost always) less than 24h. We'll often run daily standup meetings if timezones permit, but these are optional.
As an indicative guide (this is not a score sheet), in no particular order, I evaluate candidates according to:
In the past cohort I chose a diversity of candidates with varying strengths and I think this worked quite well. Some mentees were outstanding in particular dimensions, others were great all rounders.
In general I'd like projects in my stream should at least be conceptually informed by parameter decomposition, manifolds, and minimum description length framings of interpretability, if not build on them directly.
Scholars and I will discuss projects and come to a consensus on what feels like a good direction. I will not tell scholars to work on a particular direction, since, in my experience, intrinsic motivation to work on a particular direction is important for producing good research.
This stream focuses on critical challenges in AI safety and alignment, including risks from automating AI research, bottlenecks to recursive self-improvement, and the automation of safety and alignment research. Priority topics also include AGI privacy, measuring long-horizon agentic capabilities, developing new alignment methods, and advancing the science of post-training.
I usually spend at least 30 min per week in one-one-one meetings with my mentees. We can also discuss longer time slots if necessary. Besides these time slots, I try to be as responsive as possible over Slack (>2 comprehensive responses per day) and read relevant papers between weekly meetings.
I'm looking for the following skills:
I would prefer to set the overall direction, but I will listen closely to scholars about their preferences within a broad direction. Converging on a particular topic is expected to be a collaborative process.
We research early training interventions that shape a model's psychological core such that alignment generalizes through subsequent training. Potential projects span developing evaluations of model psychology, developing training interventions, experimenting with seeding the chain-of-thought patterns of the model, and methods for making models active participants in their own alignment.
By default, 1 hour weekly meetings, ideally in person in Berkeley. One day a week I will be in Berkeley and can have quite high touch in person chats throughout the day. The rest of the week I expect to be somewhat accessible on Slack and should be able to respond to messages within a few hours. I expect a lot of discussion on top of artifacts—experiment proposals, plots, slides with interim results. I can jump on quick calls to unblock, but my availability for this can change depending on the day.
Essential:
Preferred:
Not expected to be particularly helpful:
We are happy to work with fellows in the first week to jointly develop a project. We (at least Felix) are happy to be high touch during this phase and to discuss and brainstorm potential projects.
Projects on this stream cluster into a few broad areas from the empirical track: scalable oversight, AI control, monitorability and interpretability, adversarial robustness, and security.
Most fellows will work closely with one or two mentors on something that fits into the mentors' ongoing research. The above list of mentors above is tentative.
Essential:
Preferred (at least one of):
The Redwood Research stream is looking for fast empirical iterators and strategists to work on control research.
Depending on the mentor:
We are looking for people who are:
We will assign projects by default but are open to getting pitched on projects.
This stream will work on projects that empirically assess national security threats of AI misuse (CBRN terrorism and cyberattacks) and improve dangerous capability evaluations. Threat modeling applicants should have a skeptical mindset, enjoy case study work, and be strong written communicators. Eval applicants should be able and excited to help demonstrate concepts like sandbagging elicitation gaps in an AI misuse context.
The MATS Program is a 10-week research fellowship designed to train and support emerging researchers working on AI alignment, transparency and security. Fellows collaborate with world-class mentors, receive dedicated research management support, and join a vibrant community in Berkeley focused on advancing safe and reliable AI. The program provides the structure, resources, and mentorship needed to produce impactful research and launch long-term careers in AI safety.
MATS mentors are leading researchers from a broad range of AI safety, alignment, governance, field-building and security domains. They include academics, industry researchers, and independent experts who guide scholars through research projects, provide feedback, and help shape each scholar’s growth as a researcher. The mentors represent expertise in areas such as:
Key dates
Application:
The main program will then run from September 28th to December 4th, with the extension phase for accepted fellows beginning in December.
MATS accepts applicants from diverse academic and professional backgrounds - from machine learning, mathematics, and computer science to policy, economics, physics, cognitive science, biology, and public health, as well as founders, operators, and field-builders without traditional research backgrounds. The primary requirements are strong motivation to contribute to AI safety and evidence of technical aptitude, research potential, or relevant operational experience. Prior AI safety experience is helpful but not required.
Applicants submit a general application, applying to various tracks (Empirical, Theory, Strategy & Forecasting, Policy & Governance, Systems Security, Biosecurity, Founding & Field-Building.
In stage 2, applicants apply to streams within those tracks as well as completing track specific evaluations.
After a centralized review period, applicants who are advanced will then undergo additional evaluations depending on the preferences of the streams they've applied to before doing final interviews and receiving offers.
For more information on how to get into MATS, please look at this page.