实证研究

本方向涵盖通过机器学习实验开展的实践研究,旨在理解并提升模型安全性,包括 AI 控制、可解释性、可扩展监督、评估、红队测试和鲁棒性。它以研究方法而非单一研究议题为界。如果你主要运用机器学习工程方法,这个方向适合你。

‍

Application process

  • 第一阶段:完成通用申请
  • 第二阶段第一部分:完成 1 至 2 项评估研究判断力和技术实现能力的测评,并可能需要提供推荐人信息
  • 第二阶段第二部分:回答研究流方向选择问题
  • 第三阶段:参加面试和工作测试

实证研究 track overview

本方向以研究方法而非单一议题为核心。研究员通过机器学习实验,理解并提升前沿模型的安全属性,课题涉及可解释性、AI 控制、可扩展监督、评估、红队测试、鲁棒性,以及错位行为模型生物等。共同特点是通过实际模型开展工作(训练、探测、微调、测量等),而不只是从第一性原理推演。这是项目中规模最大的方向,也是进入技术 AI 安全研究最常见的途径。

我们希望研究员主要使用机器学习工程方法,且具备广义上的相关能力。核心要求是能够设计并运行针对语言模型或其他深度学习系统的实验,并根据结果快速迭代。通常,这意味着熟悉 Python(无论是否借助 AI 编程工具),了解中等规模模型运行所需的基础设施,并能判断哪些实验值得开展。使命契合度也很重要;研究员应能说明某项实证研究如何实质性降低前沿 AI 风险,而不只是它能否产出论文。与其他方向相比,学历和资历并非主要考量。以往的优秀研究员包括本科生,也包括资深行业研究人员。

我们会根据契合度为研究员匹配导师,并规划项目,使其在项目结束前产出具体成果,例如论文、评估套件、开源工具或技术报告。本方向的成果面向前沿实验室的安全与对齐团队、政府及其他评估机构,以及更广泛的机器学习研究社区。

如果你对这类研究感兴趣,欢迎申请。

实证研究 track streams

This stream focuses on secret loyalties, where an LLM covertly tries to advance a principal's interests. Secret loyalties have been established as a pressing threat [1], and model organisms of narrow secret loyalties have been constructed and audited [2]. This stream aims to advance the empirical foundations of our understanding of secret loyalties. The goal being that humanity is well-equipped to deal with secret loyalty installation attempts as and when catastrophic secret loyalties become possible in the future.

​

We think being well-equipped looks like having sufficient security measures in place in frontier AI companies, understanding the dynamics and behaviours of secretly loyal AI systems, having effective auditing and verification protocols for secret loyalties and attempts to install them, and these protocols actually being followed by relevant stakeholders.

Read more
Desired fellow characteristics

This coalition of mentors make up the “Anthropic Stream”. This stream spans a range of empirical research areas in AI safety on LLMs, including AI control, scalable oversight, model organisms, model internals, model welfare, security, and more. You’ll be pitched, and have the option to pitch, a variety of safety research projects, and then be matched to projects and mentors based on your interests/preferences on research and what you’d like to get out of MATS. Fellows in this stream frequently receive funding and continued mentorship after MATS to complete their research project, usually leading to a (co-)first author paper. People in this stream often end up in long-term homes for safety research after MATS (e.g. Anthropic, Redwood Research, OpenAI).

​

Anthropic mentors share an application, tend to collaborate and co-mentor projects together, and generally share infrastructure to streamline the fellow experience. By applying to this stream, you are being considered for all of the Anthropic mentors.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

We will continue working on black-box monitors for scheming in complex agentic settings, building on the success of the previous stream. Concretely, we will work on scaling our datasets and fine-tuning efforts, as described in the scalable monitoring agenda

​

Most likely the next projects will be about automated iterated red-team vs. blue-team games. We are currently training the blue team. We will then train the red-team and within this stream, we will try and close the loop to train them both synchronously.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

This stream focuses on building a Science of Scheming, i.e. what are the mechanisms by which future models might become schemers, even though current models are not. We want to discover empirical Scaling Trends for Scheming. For example, does deceptive alignment become easier to discover with improved model capabilities?

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

Theory of change: Soon, most important work will be done by AI. AI is going to increasingly advise people and help with important things, many of which are time-sensitive and path dependent, e.g., work on alignment/safety (including various things like how LLMs should behave given that they’re very persuasive); how to think about acausal trade; how to organize society. It seems good for AI to do well at those things.

​

Of course, a lot of the relevant skills for doing well at these tasks are the same skills that cause AI risk and that AI companies work on (and are incentivized to work on) by default; like coding, some kinds of forecasting, etc.

​

We want to make models better at things that are net positive for the future, but that likely won’t benefit much from said default training (or perhaps will even be made worse by such training – e.g., via sycophancy).

​

In practice, a lot of the tasks that we’re interested in from this perspective are what we call “conceptual”: tasks that are hard to verify and don't have clear ground truth but where we nonetheless feel like we can make progress through argument and reason.

​

You can visit conceptualreasoning.ai to get a sense of our work to date.

​

We also take a keen interest in projects directly aimed at making future acausal interactions go well.

Read more
Desired fellow characteristics

I have two broad areas.

​

Security:

I am interested in building demonstrations for hacking real-world AI deployments to show that they are not secure. The goal is to force companies to invest in alignment techniques that can solve the underlying security issues.

​

Verification:

Verification via TEEs or ZKPs

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

This stream will focus on model motivations and character, open-ended environments, and new forms of misalignment.

Read more
Desired fellow characteristics

This stream will focus on monitoring, stress-testing safety methods, and evals, with a focus on risks from scheming AIs. Examples include (black-box) AI control techniques, white-box monitors (probes etc.), chain-of-thought monitoring/faithfulness, building evaluation environments, and stress-testing mitigations.

Read more
Mentorship structure
Desired fellow characteristics
Project selection process

常见问题解答

什么是 MATS 项目?
MATS 导师是谁?
MATS 项目的关键日期有哪些?
谁有资格申请?
申请和导师选择流程是怎样的?