Sambhav Maheshwari, Jan Wehner

I work on mitigating risks from scheming or compromised AI agents deployed for high-stakes national security tasks.

Stream overview

1. **Mapping threat models for AI agent insider threats in military and intelligence contexts.** I'd like to lay out the pathways to harm that could arise from misaligned or compromised AI agents operating in high-stakes national security deployments. Some example use cases include cyber offense and defense, decision support, AI-enabled military R&D, and autonomous weapons systems.

2. **Making control sketches to mitigate AI-insider threats, focusing on AI-based monitoring and response.** I want to describe the general control architecture/approach for each use-case (e.g. where monitors sit in the workflow, what response measures are apt, what infrastructure components are needed), and expand how protocols and parameters might vary across sub-cases (e.g., network defense vs. incident response within cyberdefense; analytic triage vs. finished products within intelligence analysis).

3. **Your own project.** We're also open to mentoring fellows who want to pursue their own project in an adjacent area, assuming there's a good fit.

Mentors

Sambhav Maheshwari
Institute for AI Policy and Strategy
,
Research Associate
Washington, D.C.
No items found.

Sambhav is a research associate on the Frontier Security team at IAPS, where he focuses on AI deployments in defense and national security, standards for internally deployed models, and threat modeling.   

  

Before joining IAPS, Sambhav served as Co-director of the Cambridge AI Safety Hub, where he ran the MARS (Mentorship for Alignment Research Students) program.

Read more
Jan Wehner
Institute for AI Policy and Strategy
,
Researcher
No items found.

Jan is a researcher at the Institute for AI Policy and Strategy (IAPS), where he works on technical AI governance to reduce catastrophic risk from AI. He is currently threat modelling how national security uses of frontier AI could go really badly and developing honeypots to secure internal AI agents.

Before joining IAPS, he was a GovAI winter fellow, a Pivotal fellow and a PhD candidate at the CISPA Helmholtz Center for Information Security. His past research spans Interpretability, ML security and AI Alignment. He holds a BA in Information Systems and a MSc in CS.

Read more

Mentorship style

Fellows we are looking for

1. **Essential**:

    

  **Strong writing:** Can write clearly and precisely for both technical and policy audiences (or has shown clear potential to).

  **Transparent reasoning:** Excels at conceptual reasoning, lays out arguments and assumptions explicitly, and brings a skeptical mindset.  

  **Research Independence:**  Has previously taken ownership of a research project and can make progress without frequent direction on what to do next.   

  **GCR Familiarity:** Is broadly familiar with the core GCR/TAIS literature, and if not, has demonstrated ability to upskill quickly on relevant topics.

Project selection

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Control, Monitoring, Red-Teaming, Scalable Oversight, Scheming & Deception
SF Bay Area
Agent Foundations
SF Bay Area
Interpretability
SF Bay Area
Interpretability, Monitoring, Dangerous Capability Evals
SF Bay Area
Interpretability, Model Organisms, Red-Teaming, Safeguards, Scheming & Deception
SF Bay Area
Interpretability
Grand Rapids
Agent Foundations
Washington, D.C.
Compute Infrastructure, Policy & Governance, Security
London
Dangerous Capability Evals, Compute Infrastructure, Policy & Governance, Strategy & Forecasting
Washington, D.C.
Compute Infrastructure, Security
London
Control, Monitoring
London
Control, Scheming & Deception, Dangerous Capability Evals, Model Organisms, Monitoring
SF Bay Area
Security, Dangerous Capability Evals
SF Bay Area
Dangerous Capability Evals
Boston
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure