Conceptual research on deceptive alignment, designing scheming propensity evaluations and honeypots. Some example directions:
London
I am a research scientist on the AGI Safety & Alignment team at Google DeepMind. I focus on deceptive alignment and AI control, particularly [scheming propensity evaluations](https://arxiv.org/abs/2605.29729). My past research includes dangerous capability evals, power-seeking incentives, specification gaming, and avoiding side effects.
I will talk through project ideas with scholars