I research AI safety and alignment at Anthropic. Before that, I was a research scientist at Google DeepMind. I completed my PhD at UC Berkeley's Center for Human-Compatible AI, advised by Stuart Russell. I previously cofounded FAR.AI, a 501(c)3 research nonprofit that incubates and accelerates beneficial AI research agendas.
I develop AI alignment frameworks, stress-test their limits, and turn insights into methodology adopted across the field. I have established that chain-of-thought monitoring is a substantial defense when reasoning is necessary for misalignment, designed practical metrics to preserve monitorability during model development, shown that obfuscated activations can bypass latent-space defenses, and developed StrongREJECT, a jailbreak benchmark now used by OpenAI, US/UK AISI, Amazon, and others.
Krishnamurthy (Dj) Dvijotham is a senior staff research scientist at Google DeepMind, where he leads efforts on the development of secure and trustworthy AI agents. He previously founded the AI security research team at ServiceNow Research and co-founded the robust and verified AI team at DeepMind. His past research has received best paper awards at many leading AI conferences, including most recently at ICML and CVPR 2024. His research led to the framework used for AI security testing at ServiceNow and has been deployed in several Google products, including the Android Play Store, YouTube and Gemini.
I am an interpretability researcher at Anthropic. I am most interested in simple, practical interpretability approaches that are targeted at making models safer. In a previous life, I worked as a neuroscientist.
I work at the [UK AI Security Institute](https://www.aisi.gov.uk/). In the past, I’ve done research in high-performance computing, language model pretraining, interpretability, and [hardware enabled governance](https://flexheg.com/).
Seth is a senior advisor on research and policy at SecureBio, where he works on AI and biosecurity. He previously completed a Ph.D. at Harvard and postdoctoral research at the University of Chicago.
Shi Feng leads a research group working on oversight and control. He is an assistant professor at George Washington University. Prior to that, he was a postdoc in the NYU Alignment Research Group under Sam Bowman. He currently focuses on deception and collusion, with an emphasis on propensity and evaluation realism.
Adam Shai has extensive research experience in experimental and computational neuroscience. He earned his PhD from Caltech and has over a decade of experience investigating the neural basis of intelligent behavior, most recently as a researcher at Stanford. Driven by the pressing need for AI safety, he has now turned his expertise to neural networks, aiming to develop principled methods for controlling and aligning increasingly advanced AI systems.
Adam co-founded and now leads research at Simplex, an organization dedicated to building a science of representations in AI systems.
I was until recently a professional mathematician at the University of Melbourne, where I worked on algebraic geometry, mathematical logic, some aspects of mathematical physics, and most recently statistical learning theory. As of early 2025 I left academia to direct research at Timaeus on AI safety.
Saad Siddiqui is a senior researcher at Safe AI Forum, where his research examines possible agreement between leading AI powers. He previously worked as a management consultant at Bain and Company in Singapore.
Oly works at the Future of Life Foundation on sourcing and developing ambitious ideas to build a flourishing future, grounded in realistic scenarios for AI and other technological development. Priorities include human collective intelligence uplift, gentle and manageable multiagent transitions, and defensive tech.
Oly previously worked on loss of control risk modelling and evaluation at the UK AI Safety/Security Institute and continues to engage with the OECD, UK FCDO, DSIT, and parliamentarians on AI governance.
He researched (LM) agent oversight and multiagent safety at Oxford and was one of the first beneficiaries of the MATS program in 2021-22. Before his AI safety work, he was a senior data scientist and software engineer.
MATS 项目是一项为期 10 周的研究奖学金计划,旨在培养和支持从事人工智能对齐、透明度和安全领域工作的新兴研究人员。研究员将与世界一流的导师合作,获得专门的研究管理支持,并加入位于伯克利、致力于推动人工智能安全与可靠发展的活跃社区。该项目提供开展高影响力研究并开启人工智能安全领域长期职业生涯所需的架构、资源和指导。
MATS 导师均为来自人工智能安全、对齐、治理、领域建设及安全等广泛领域的顶尖研究人员。他们包括学术界人士、行业研究员以及独立专家,负责指导学者开展研究项目、提供反馈,并助力每位学者的研究成长。导师们的专业领域涵盖:
查看 往届及现任导师
关键日期
申请:
主项目将于 9 月 28 日至 12 月 4 日进行,获选研究员的延展阶段将于 12 月开始。
MATS 欢迎来自不同学术和专业背景的申请者——从机器学习、数学和计算机科学,到政策、经济学、物理学、认知科学、生物学和公共卫生,同时也欢迎没有传统研究背景的创业者、运营人员和领域建设者。主要要求是具备为人工智能安全做出贡献的强烈动机,并展现出技术能力、研究潜力或相关的运营经验。具备人工智能安全相关经验会有所帮助,但并非必要条件。