Daniel is a professor of computer science at UIUC, where he studies the progress of AI, with a particular focus on dangerous capabilities of AI agents. His work includes:
- CVE-Bench, an award winning benchmark (SafeBench award, ICML spotlight) that is used by frontier labs and governments to measure AI agents' ability to find and exploit real-world vulnerabilities.
- Agent Benchmark Checklist, an award winning work (Berkeley AI summit, 1st place Benchmarks & Evaluations track) that highlights major issues in existing benchmarks.
- InjecAgent, one of the first AI agent safety benchmarks, used by governments and major labs.
Sara Price leads Alignment Training at Anthropic. Sara's research focuses on alignment training that generalizes reliably, including methods for reducing agentic misalignment.
Roger Grosse is an associate professor of computer science at the University of Toronto and a member of Anthropic’s Alignment Science team, where he works on training data attribution. He earned his PhD in computer science from MIT.
Xander Davies is a Member of the Technical Staff at the UK AI Security Institute, where he leads the Red Teaming group, which uses adversarial ML techniques to understand, attack, and mitigate frontier AI safeguards. He is also a PhD student at the University of Oxford, supervised by Dr. Yarin Gal. He previously studied computer science at Harvard, where he founded and led the Harvard AI Safety Team.
I am a principal investigator at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, where I lead the AI Safety and Alignment group. I also serve as chapter lead for the new edition of the International AI Safety Report chaired by Prof. Yoshua Bengio. I have worked on AI safety with leading organizations in the field (OpenAI, Anthropic, UK AI Safety Institute, Center for AI Safety, Gray Swan AI). I obtained my PhD in machine learning from EPFL in 2024 advised by Prof. Nicolas Flammarion. My PhD thesis was awarded the Patrick Denantes Memorial Prize for the best thesis in the CS department of EPFL and was supported by the Google and Open Phil AI PhD Fellowships.
I’m a Research Scientist in MIT CSAIL with the MIT-IBM Watson AI Lab. I did my PhD in Brain and Cognitive Sciences at MIT, as an NSF Fellow working with Josh Tenenbaum and Antonio Torralba. My work investigates representations underlying intelligence in artificial (and previously, biological) neural networks.
Jacob Hilton is a researcher at the Alignment Research Center (ARC), a nonprofit working on the theoretical foundations of mechanistic interpretability. He previously worked at OpenAI on reinforcement learning from human feedback, scaling laws and interpretability. His background is in pure mathematics, and he holds a PhD in set theory from the University of Leeds, UK.
Eli is working on AI scenario forecasting with the AI Futures Project, where he co-authored AI 2027. He advises Sage, an organization he cofounded that works on AI Digest (interactive AI explainers) and forecasting tools. He previously worked on the AI-powered research assistant Elicit.
As AI Science Advisor to the California Governor’s Office of Emergency Services (Cal OES), Michael Chen advises senior leadership on frontier AI safety and risk assessment, with a particular emphasis on critical safety incidents, AI and cyber defense, and risk from developers’ internal deployment of AI, such as sabotage by AI agents and automated AI R&D. As Science Advisor, he coordinates with AI governance leads across California’s state government and facilitates cross-sector collaboration with the academic research community, the private sector, and community and nonprofit organizations.
Michael previously worked on evaluations-based AI governance at METR, an independent California-based nonprofit evaluator of autonomous AI agent capabilities and risks. He advised leading AI developers on frameworks for assessing, mitigating, and transparently disclosing catastrophic AI risks. He also assisted with third-party evaluations, including a review of a developer’s report assessing sabotage risk from AI agents, and contributed to a catalog of incidents in which AI agents acted beyond their operators’ intent. Michael has engaged with U.S. government bodies on frontier AI evaluation, including the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST), and conducted research at UC Berkeley’s Center for Human-Compatible AI on learning human preferences for language models. His research and commentary on AI have been covered in outlets such as Time, The Guardian, and MIT Technology Review. Michael is a part-time PhD student at the University of Oxford and an affiliate of the Oxford Martin AI Governance Initiative.
I’m a Member of Technical Staff at OpenAI working on monitoring LLM agents for misalignment. Previously, I worked on AI control and safety cases at the UK AI Security Institute and on honesty post-training at Anthropic. Before that, I did a PhD at the University of Sussex with Chris Buckley and Anil Seth focusing on RL from human feedback (RLHF) and spent time as a visiting researcher at NYU working with Ethan Perez, Sam Bowman and Kyunghyun Cho.
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。