Sambhav Maheshwari, Jan Wehner

I work on mitigating risks from scheming or compromised AI agents deployed for high-stakes national security tasks.

Stream overview

1. Mapping threat models for AI agent insider threats in military and intelligence contexts. I'd like to lay out the pathways to harm that could arise from misaligned or compromised AI agents operating in high-stakes national security deployments. Some example use cases include cyber offense and defense, decision support, AI-enabled military R&D, and autonomous weapons systems.

2. Making control sketches to mitigate AI-insider threats, focusing on AI-based monitoring and response. I want to describe the general control architecture/approach for each use-case (e.g. where monitors sit in the workflow, what response measures are apt, what infrastructure components are needed), and expand how protocols and parameters might vary across sub-cases (e.g., network defense vs. incident response within cyberdefense; analytic triage vs. finished products within intelligence analysis).

3. Your own project. We're also open to mentoring fellows who want to pursue their own project in an adjacent area, assuming there's a good fit.

Mentorship style:

Standard (1-2 hours of weekly 1:1s)

Location during program:

Washington, D.C.

London location preference:

Strong preference

Berkeley location preference:

Required

Mentors

Sambhav Maheshwari
Institute for AI Policy and Strategy
,
Research Associate
No items found.

Sambhav is a research associate on the Frontier Security team at IAPS, where he focuses on AI deployments in defense and national security, standards for internally deployed models, and threat modeling.   

  

Before joining IAPS, Sambhav served as Co-director of the Cambridge AI Safety Hub, where he ran the MARS (Mentorship for Alignment Research Students) program.

Read more
Jan Wehner
Institute for AI Policy and Strategy
,
Researcher
No items found.

Jan is a researcher at the Institute for AI Policy and Strategy (IAPS), where he works on technical AI governance to reduce catastrophic risk from AI. He is currently threat modelling how national security uses of frontier AI could go really badly and developing honeypots to secure internal AI agents.

​

Before joining IAPS, he was a GovAI winter fellow, a Pivotal fellow and a PhD candidate at the CISPA Helmholtz Center for Information Security. His past research spans Interpretability, ML security and AI Alignment. He holds a BA in Information Systems and a MSc in CS.

Read more

Fellows we are looking for

1. Essential:

​

Strong writing: Can write clearly and precisely for both technical and policy audiences (or has shown clear potential to).

Transparent reasoning: Excels at conceptual reasoning, lays out arguments and assumptions explicitly, and brings a skeptical mindset.

Research Independence: Has previously taken ownership of a research project and can make progress without frequent direction on what to do next.

GCR Familiarity: Is broadly familiar with the core GCR/TAIS literature, and if not, has demonstrated ability to upskill quickly on relevant topics.

Any of the following would strengthen the application:

​

1. Military/intelligence background: Has been in, or spent substantial time working on topics relating to, the military or IC.

2. Technical AI safety background: Has conducted research in technical AI safety, and in particular, AI control.

3. LLM power user: Is skilled at getting useful conceptual work out of LLMs.

4. Policy background: Has written memos or policy reports for a non-technical audience, and/or thought extensively about AI governance.

Project selection