Conceptual research on deceptive alignment, designing realistic scheming propensity evaluations and honeypots.
Standard (1-2 hours of weekly 1:1s)
London
Required
Indifferent
I am a research scientist on the AGI Safety & Alignment team at Google DeepMind. I focus on deceptive alignment and AI control, particularly scheming propensity evaluations. My past research includes dangerous capability evals, power-seeking incentives, specification gaming, and avoiding side effects.
Combination of strong conceptual research ability, engineering ability, and AI alignment knowledge.
I will talk through project ideas with scholars