本方向研究用于监控和保护 AI 开发与部署的软件及基础设施层安全机制,包括侧信道分析、集群安全和物理层验证。本方向所说的“系统安全”不同于通常指对抗鲁棒性或越狱的广义“AI 安全”;重点是保护先进 AI 所依赖的系统,包括数据中心、硬件供应链、算力集群和模型权重。
本方向关注前沿 AI 系统运行所依赖的硬件、软件和基础设施。即使对齐研究取得成功,只有在模型权重不会被盗、训练与推理算力不会遭到篡改、运行先进 AI 的系统能够接受审计与验证时,其安全属性才有实际意义。本方向旨在为此建立技术基础,研究流涵盖模型权重保护、侧信道分析、安全飞地、硬件供应链保障、数据中心与集群安全、算力使用的物理层验证,以及让算力治理和出口管制得以执行的技术组件。
我们希望研究员在系统或安全领域有扎实的专业能力,并熟悉至少一个相关领域:安全工程、漏洞研究、密码学、操作系统等底层软件、硬件安全(例如芯片、FPGA 和嵌入式系统),或基础设施层的系统工程。机器学习经验是加分项,但不是必需条件。关键在于能够审慎分析攻击者、侧信道和信任边界。以往的优秀申请者来自科技公司的安全团队、硬件与芯片设计、密码学研究、政府安全工作、CTF 与漏洞研究,以及嵌入式系统工程等领域。
我们会根据契合度为研究员匹配导师,并规划项目,使其在项目结束前产出具体成果,例如安全审计、技术规范、防御原型、攻击演示或验证协议提案。本方向的成果面向实验室安全与基础设施团队、硬件供应商,以及需要技术基础来执行算力治理和出口管制的政策相关方。
I have two broad areas.
Security:
I am interested in building demonstrations for hacking real-world AI deployments to show that they are not secure. The goal is to force companies to invest in alignment techniques that can solve the underlying security issues.
Verification:
Verification via TEEs or ZKPs
I will meet 1-1 or as a group, depending on the interests as they relate to the projects. Slack communication outside of the 1-1.
I strongly prefer multiple short meetings over single long meetings, except at the start.
I'll help with research obstacles, including outside of meetings
For security:
You should have a strong security mindset, having demonstrated the willingness to be creative on this. I would like to see past demonstration of willingness to get your hands dirty and try many different systems.
For benchmarks:
As creative as possible, willingness to work on the nitty gritty, willingness to work really hard on problems other people find boring. Interests as far away from SF-related interests as possible.
Mentor(s) will talk through project ideas with scholar
In this project, we will explore GPU side-channel attacks to extract information about model usage. A simple example is to observe (via radio, power fluctuations, acoustics, etc.) which experts were used in each forward pass of an MOE model, then use those observations to guess which tokens were produced.
Co-working 2-4 hours per week, including detailed guidance. Flexible. 1 hour check-ins per week. You can schedule ad-hoc calls if stuck or wanting to brainstorm.
Please note: experience with hardware is not a requirement for this stream, as long as you are willing to work hard and learn fast, and can show other evidence of exceptional ability. If in doubt: we encourage you to apply!
We will provide you with a lot of autonomy and plug-and-play access to a rare combination of tools and equipment—in exchange we expect you to have a strong self-direction, intellectual ambition, and a lot of curiosity. This stream requires you to have a tight experiment loop to form and test hypotheses on the fly.
Example skill profiles:
Must have: Trained or fine-tuned a transformer language model in PyTorch (toy models and following guides is fine). Familiar with basic electronics concepts (voltage, current, transistors). Has experience writing research papers, even as a class assignment.
Nice to have: Familiarity with LaTeX, PyTorch internals, CUDA/OpenCL, GPU architecture, chip design, oscilloscopes, signal processing, electrical engineering.
There is a cluster of potential projects to choose from. As a team, we will decide which to pursue based on individual interest and skills. Mentors will pitch example projects and scholars can then modify and re-pitch them. Once the research problem, hypothesis, and testing plan are written and agreed on, scholars begin object-level work. We encourage failing fast and jumping to a fallback project.
The SL5 Task Force will build out a prototype SL5 datacenter this year together with frontier AI labs. This will be a massive research and engineering project with many avenues for spinning out new organizations and research programs. This project is urgent due to this technology being needed in the next 1 to 2 years.
We will meet at least 1h a week synchronously and communicate daily via stand-ups on slack. I typically respond within a few hours for additional feedback and within 1-3 days for indepth code or other review. Scholars can also schedule adhoc calls with me or my co-mentor Luis if they're stuck.
You may have the option to join company meetings and work from our offices 1+ days a week to collaborate with SL5 engineering staff.
This stream is best for strong technical IC's looking to move into research lead / tech lead / org lead positions in the future.
Essential:
Preferred:
Not a good fit:
Our stream focuses on AI verification, as in how actors can check that the use of AI compute is compliant with policy, especially for enabling international agreements on AI. This sense of verification is much broader than formal verification.
We'll meet once or twice a week (~1 hr/wk total, as a team if it's a team project). I'm based in DC, so we'll meet remotely. I (Mauricio) will also be available for async discussion, career advising, and detailed feedback on research plans and drafts.
Strong analytical and writing skills, research pragmatism and judgment, fast learner, proactive, and AI landscape context.
I'll talk through project ideas with scholar
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。