Francis Rhys Ward

Currently I'm focusing on projects which evaluate sabotage, sandbagging, and oversight subversion in frontier agents (aka, building control evals). Additionally, I'm interested in coming up with novel alignment or control plans.

Stream overview

Two concrete project proposals:

  1. Realistic benchmarks for evaluating oversight subversion
    1. A core bottleneck to scalable/automated oversight research is high-quality, realistic benchmarks for studying oversight subversion. To help build effective oversight protocols, this project aims to evaluate AI agents’ capabilities to autonomously subvert oversight mechanisms in realistic environments.
  2. Collusion in automated oversight regimes
    1. Past work typically assumes that the automated overseer is trustworthy, for example, because it is not capable enough to pose a threat. We aim to relax this assumption to study the problem that frontier AI systems might collude with one-another to subvert human control.

Mentors

Francis Rhys Ward (Rhys)
Arrow Research
,
Director
London
No items found.

Rhys is a researcher at LawZero. He recently finished his PhD, on formalising and evaluating AI deception.

Technically, his work involves both conceptual research, in the intersection of game theory, causality, and philosophy, in addition to empirical evaluations of frontier AI systems. Rhys is now focusing on control-style research and frontier LM agent alignment.

He is a member of Tom Everitt's Causal Incentives Working Group. Previously, Rhys has worked at the Centre for Assuring Autonomy, the Center on Long-Term Risk, the Centre for the Governance of AI, and the UK’s AI Safety Institute.

In this MATS cohort, he is looking to work on evals and control-style research.

Read more

Mentorship style

Fellows we are looking for

  • Good communication -- it's important that you keep me updated with how you're doing and how I can help!
  • Experience with LM experiments (e.g., fine-tuning).
  • Clear thinker and writer.

Project selection

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Montreal
Empirical
SF Bay Area
Empirical
Theory
London
Empirical
Founding and Field-Building
Policy and Governance
SF Bay Area
Empirical
Systems Security
London
Empirical
Washington, D.C.
Policy and Governance
Washington, D.C.
Policy and Governance
SF Bay Area
Strategy and Forecasting
Policy and Governance
SF Bay Area
Founding and Field-Building
Washington, D.C.
Biosecurity
Empirical
London
Empirical
London
Empirical