Sunishchal Dev

I work on the science of evaluating advanced AI systems for biological and CBRN risks, with a particular interest in translating technical evidence into decisions by governments and frontier AI developers. In this stream, I’m interested in developing novel capability evaluations, studying how dangerous or dual-use capabilities diffuse into increasingly accessible models, and building scalable red-teaming methods that produce rigorous, decision-relevant evidence without requiring risky real-world demonstrations.

Stream overview

 I expect to refine project selection with fellows at the beginning of the program based on their technical backgrounds, emerging research opportunities, and which questions seem most likely to produce decision-relevant results. Example projects could include:

  1. Measuring the diffusion of bio-relevant capabilities into open-weight models. A fellow could develop methods for tracking when capabilities first demonstrated by frontier closed models become accessible in open-weight systems, and study how factors such as model scale, post-training, scaffolding, tool use, or elicitation affect that diffusion. The goal would be to develop better indicators of when potentially consequential capabilities are becoming broadly accessible.
  2. Developing evaluations for operationally relevant biological capabilities. A fellow could build and validate novel benchmarks or agentic evaluations that target capabilities poorly captured by existing static benchmarks. For example: scientific troubleshooting, multi-step research workflows, or the ability to substitute for scarce expert knowledge. I am particularly interested in evaluations that help distinguish impressive benchmark performance from capabilities that would actually provide substantial utility to real practitioners.
  3. Automated red-teaming and evaluation methods for CBRN risk. A fellow could develop scalable methods for discovering concerning model behaviors, stress-testing safeguards, or systematically eliciting latent capabilities across different models and deployment settings. This could include studying how sensitive risk assessments are to prompting, scaffolding, tool access, or other forms of capability elicitation, while keeping experiments safe and primarily in silico.
Mentorship style:

Standard (1-2 hours of weekly 1:1s)

Location during program:

Washington, D.C.

London location preference:

Weak preference

Berkeley location preference:

Strong preference

Mentors

Sunishchal Dev (Dev)
Center for AI Standards and Innovation; RAND
,
AI Evaluations Research Scientist
Biosecurity
Capability and Propensity Evaluations

Sunishchal Dev is an AI safety researcher working pre deployment testing with a focus on biosecurity at the U.S. AI Safety Institute within NIST’s Center for AI Standards and Innovation (CAISI). Previously, he was an AI Evaluations Research Scientist at RAND, where he led machine learning engineering efforts focused on evaluating frontier AI systems, including building biological capability benchmarks, assessing risks from open-weight models, and developing methods to make LLM-based evaluations more reliable. He was a fellow during MATS 6.0 under the mentorship of Marius Hobbhahn. Before moving into AI safety, Dev spent nearly a decade as a data scientist, machine learning engineer, and management consultant.

Read more

Fellows we are looking for

  • Strong technical implementation skills. I’m looking for fellows who are comfortable writing Python, working with modern language models and APIs, building experimental pipelines independently, and testing/debugging their code.
  • Demonstrated ability to conduct open-ended research. You should have completed at least one substantial research or technical project where the answer was not known in advance and you had meaningful ownership over the approach, execution, or interpretation.
  • Ability to reason quantitatively about empirical evidence. My projects will involve designing evaluations, comparing models, analyzing noisy results, and deciding what conclusions the evidence does and does not support. You should be comfortable with basic statistics and careful experimental design.
  • Ability and willingness to engage deeply with biology. A formal biology degree is not required, but you should either already have some biological knowledge or be able to rapidly develop enough domain understanding to reason critically about biological tasks, scientific workflows, and the validity of biosecurity evaluations.
  • Strong research judgment and safety awareness. Some projects may involve dual-use biological or CBRN questions. I am looking for fellows who can distinguish scientifically useful experiments from unnecessarily hazardous demonstrations and who are comfortable modifying a research plan when the risk-benefit tradeoff is poor.
  • Ability to work independently and communicate clearly. I expect fellows to drive much of the day-to-day research themselves, identify blockers early, and explain technical results clearly in writing and discussion.

Project selection