UKAISI Red-Team

This stream is for the UK AISI Red-team. The team focuses on stress-testing mitigations for AI risk, including misuse safeguards, control techniques and model alignment red-teaming. We plan to work on projects building and improving methods for performing these kinds of evaluations and methods.

Stream overview

Project could include:

  • developing methods for automated red-teaming of mitigations for any of the risk areas
  • building evaluation environments for evaluating misuse of agentic AI systems
  • developing attacks against asynchronous or stateful monitoring systems for misuse
  • developing and red-teaming control protocols and modelling effect in realistic deployment situations
  • improving automated auditing tools such as Petri and Bloom to be more realistic, controllable, and with additional beneficial affordances for propensity evaluations; building propensity evaluations within automated auditing tools or otherwise.

Location during program:

London

Mentors

Xander Davies
UK AISI, University of Oxford
,
Safeguards Team Lead
Capability and Propensity Evaluations
AI 系统安全
对抗鲁棒性与安全防护

Xander Davies is a Member of the Technical Staff at the UK AI Security Institute, where he leads the Red Teaming group, which uses adversarial ML techniques to understand, attack, and mitigate frontier AI safeguards. He is also a PhD student at the University of Oxford, supervised by Dr. Yarin Gal. He previously studied computer science at Harvard, where he founded and led the Harvard AI Safety Team.

Read more
Robert Kirk
UK AISI
,
Research Scientist
Misalignment Science
Capability and Propensity Evaluations
对抗鲁棒性与安全防护

Robert is a research scientist and the acting lead of the alignment red-teaming sub-team at UK AISI. This team's focus is on stress-testing model alignment to detect and understand model propensities relevant to loss-of-control risks. Before that, he's most recently worked on misuse research, focusing on evaluations of safeguards against misuse and mitigations for misuse risk, particularly in open-weight systems. He graduated from his PhD from University College London on generalisation in LLM fine-tuning and RL agents in January 2025.

Read more
Alexandra Souly (Alex)
UK AISI
,
Technical Staff
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations
AI 系统安全
对抗鲁棒性与安全防护

Alex Souly is a researcher on the Red Team at the UK AI Security Institute, where she works on the safety and security of frontier LLMs. She has contributed to pre-deployment evaluations and red-teaming of misuse safeguards and alignment (see Anthropic and OpenAI blogpost), and worked on open source evals like StrongReject and AgentHarm. Previously, she studied Maths at Cambridge and Machine Learning at UCL as part of UCL Dark lab, interned at CHAI, and in another life worked as a SWE at Microsoft.

Read more
Eric Winsor
UK AISI
,
Research Engineer
AI Control and Monitoring
Capability and Propensity Evaluations
对抗鲁棒性与安全防护

Eric Winsor is a research scientist at the UK AI Security Institute and contributes to adversarial testing of frontier AI model safeguards. Winsor earned a B.S.E. in computer engineering from the University of Michigan.

Read more
Giorgi Giglemiani
UK AISI
,
Research Engineer
Interpretability
Capability and Propensity Evaluations
对抗鲁棒性与安全防护

Giorgi Giglemiani works at the UK AI Security Institute and coauthored Boundary Point Jailbreaking. Previously, Giglemiani researched synthetic activations composed of sparse-autoencoder latents at LASR Labs.

Read more
Asa Cooper Stickland
UK AISI
,
Research Scientist
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations
对抗鲁棒性与安全防护

I'm a research scientist at the UK AI Security Institute, working on AI control red teaming and model organisms of misalignment. I was previously a postdoc with Sam Bowman at NYU, did MATS with Owain Evans, and mentored for the MATS, SPAR and Pivotal fellowships. I got my PhD at the University of Edinburgh, supervised by Iain Murray.

Read more

Fellows we are looking for

We're looking for scholars with hands-on experience in machine learning and AI security, particularly those interested in adversarial robustness, red teaming, or AI safeguards. Ideal candidates would have:

​

  • Experience with large language models (training, fine-tuning, evaluation, or safety research)
  • Strong technical foundations in ML, ideally with coding experience in PyTorch or Inspect
  • Interest in one or more of our three focus areas: misuse (securing systems against bad actors), alignment (ensuring AI systems behave as intended), or control (keeping AI systems under human control even when misaligned)
  • A mission-driven mindset and curiosity about how AI security research can inform real-world policy and deployment decisions
  • An ability to advocate for your own research ideas and work in a self-directed way, while also collaborating effectively and prioritizing team efforts over extensive solo work.

​

We welcome scholars at various career stages especially those who are eager to work on problems with direct impact on how frontier AI is governed and deployed.

Project selection

Scholars will choose from a set of predefined project directions aligned with our current research priorities, such as:

  • Developing automated methods to test AI misuse safeguards
  • Investigating data poisoning attacks and defenses
  • Designing benchmarks for misuse detection across multiple model interactions
  • Testing control measures for potentially misaligned AI systems

​

We'll provide initial direction and guidance on project scoping, then scholars will have autonomy to explore specific approaches within that framework.

Expect weekly touchpoints to ensure progress and refine directions.

If mentees have particular ideas they're excited about that they see as fitting within the scope of the team's work, they're welcome to propose them, but there is no guarantee they will be selected