Francis Rhys Ward

I'm open to quite a broad range of work. By default I expect scholars to work on reasonably conceptually-simple empirical projects, such as building capability or propensity evals.

Stream overview

Alignment-science style work eg:

  1. Train model organisms of scheming.
  2. Study how character training influences reward hacking rates and how models reason about reward hacking.

Tech-gov style work:

  1. Write a report detailing public information on frontier AI company monitoring and control practices.
  2. Write a report detailing public information on training episode 'length' of frontier AIs.

Other:

  1. Build agentic propensity evaluations for reward hacking, apparent-success seeking, grader-approval, etc.
  2. Work on measuring AI R&D speed-up / indicators of RSI.
  3. Conceptual work, e.g., criticising existing threat models, or on the deal-making agenda.

Mentors

Francis Rhys Ward (Rhys)
Arrow Research
,
Director
London
No items found.

Rhys is the research director of Arrow: a new AI safety non-profit based in London.

He is currently most excited about working on alignment science (e.g., model organisms research) and technical governance style work which reduces public uncertainty regarding important questions (e.g., his recent work on no-CoT time-horizons). 

Rhys started working on AGI risk in 2019. He has previously worked at: Redwood Research, LawZero, UK AISI, GovAI, CLR, and the Centre for Assuring Autonomy. He did his PhD in AI Deception at Imperial College London.

Read more

Mentorship style

Fellows we are looking for

Technical skills, experience building evals with inspect, or working on agents or fine-tuning models with tinker. 

It's good if you have at least one completed AI safety project.

Project selection

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception