The team's focus is on stress-testing model alignment to detect and understand model propensities relevant to loss-of-control risks, and includes work on building realistic alignment evaluations, measuring and mitigating evaluation awareness, inferring hidden propensities and developing automated algorithms to search for misalignment behaviour. We perform pre-deployment alignment testing across frontier AI companies. See e.g.:
We plan to pitch specific projects at the start of the program. Below is a list of example projects which we've pitched previously (some of which we may pitch again). More detailed descriptions of example projects can be found here: https://docs.google.com/document/d/16QFQ1pdEmJ299_4HRWTHddeVxKu5VRxPgxpbDgCCFB4/edit?tab=t.0. Note that some of these have already been picked up by current mentees, but they give a sense of the kind of work we're interested in.
Current models have a propensity to mis- or over-represent their work on difficult, long-horizon, open-ended tasks. As described in more detail in this post, this could lead to catastrophic outcomes if not fixed, and if AIs are inappropriately deferred to for performing AI safety research.
This project would build an evaluation to measure this behaviour, and in doing so build methodology and tooling to enable this kind of evaluation to be scalably iterated on for future models. For a given task difficulty, models are likely to misrepresent less as they get more capable (since they can perform the task correctly, and so can accurately represent success); but for tasks at the frontier of difficulty this may remain a problem. I expect that methodology built around measuring discrepancies between a model's behaviour and its final stated outputs — using grading compute scalably, inspired by scalable oversight methods — will be able to make progress on measurement here.
Alongside building evaluations targeted at specific model propensities (as above), it is important to develop scalable methods for finding arbitrary misbehaviour in models. This direction has received less attention than it should have, in my view. I expect that in the future misaligned behaviour will be hard to find, and so to be confident we have stress-tested models sufficiently we will need strong and adaptive red-teaming evaluations, rather than just static ones.
This project would develop methods for automated misalignment search, starting from existing work (Transluce, pathological behaviors and arXiv:2602.05910), and exploring evolutionary search techniques for finding contexts that induce misaligned behaviour.
This project would investigate the relationships between different methods for measuring evaluation awareness (prompting, chain-of-thought scanning, and various white-box methods), and — for models that demonstrate evaluation-gaming — how that behaviour manifests across each of these measurements. In particular, it would measure the extent to which current models can and do evaluation-game with unverbalised evaluation awareness. This would help us understand how much evidence propensity evaluations give, even in the absence of explicit evaluation awareness.
This project could also investigate methods for mitigating evaluation awareness, either through model-level interventions (e.g. steering) or through methods for automatically improving environment realism. We have upcoming work showing promise with simple methods for scalably improving environment realism, and I expect there to be many potential follow-ups to this.
Alex Souly is a researcher on the Red Team at the UK AI Security Institute, where she works on the safety and security of frontier LLMs. She has contributed to pre-deployment evaluations and red-teaming of misuse safeguards and alignment (see Anthropic and OpenAI blogpost), and worked on open source evals like StrongReject and AgentHarm. Previously, she studied Maths at Cambridge and Machine Learning at UCL as part of UCL Dark lab, interned at CHAI, and in another life worked as a SWE at Microsoft.
Robert is a research scientist and the acting lead of the alignment red-teaming sub-team at UK AISI. This team's focus is on stress-testing model alignment to detect and understand model propensities relevant to loss-of-control risks. Before that, he's most recently worked on misuse research, focusing on evaluations of safeguards against misuse and mitigations for misuse risk, particularly in open-weight systems. He graduated from his PhD from University College London on generalisation in LLM fine-tuning and RL agents in January 2025.
Each scholar will have one primary mentor from the Red Team who will provide weekly guidance and day-to-day support
Scholars will also have access to secondary advisors within their specific sub-team (misuse, alignment, or control) for technical deep-dives
Team lead Xander Davies and advisors Geoffrey Irving and Yarin Gal will provide periodic feedback through team meetings and project reviews
For scholars working on cross-cutting projects, we can arrange mentorship from multiple sub-teams as needed
Structure:
Weekly 1:1 meetings (60 minutes) with primary mentor for project updates, technical guidance, and problem-solving
Asynchronous communication via Slack/email throughout the week for quick questions and feedback
Bi-weekly team meetings where scholars can present work-in-progress and get broader team input
Working style:
We expect scholars to work semi-independently – taking initiative on their research direction while leveraging mentors for guidance on technical challenges, research strategy, and navigating AISI resources
Scholars will have access to our compute resources and operational support to focus on research
We encourage scholars to document their work and, if appropriate, aim for publication or public blog posts
We're looking for scholars with hands-on experience in machine learning and AI security, particularly those interested in adversarial robustness, red teaming, or AI safeguards. Candidates should have:
Scholars will choose from a set of predefined project directions aligned with our current research priorities, such as:
We'll provide initial direction and guidance on project scoping, then scholars will have autonomy to explore specific approaches within that framework.
Expect weekly touchpoints to ensure progress and refine directions.
If mentees have particular ideas they're excited about that they see as fitting within the scope of the team's work, they're welcome to propose them, but there is no guarantee they will be selected
The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.