Truthful AI

This stream will focus on evaluating dangerous capabilities in language models and detecting deception and dishonesty.

Stream overview

  • Defining and evaluating situational awareness in LLMs (relevant paper)
  • Predicting the emergence of other dangerous capabilities in LLMs (e.g. deception, agency, misaligned goals)
  • Studying emergent reasoning at training time (“out-of-context” reasoning). See Reversal Curse.
  • Detecting deception and dishonesty in LLMs using black-box methods
  • Enhancing human epistemic abilities using LLMs (e.g., AutocastTruthfulQA)

Location during program:

SF Bay Area

Mentors

Owain Evans
Truthful AI
,
Research Lead
Misalignment Science
Capability and Propensity Evaluations

Owain has a broad interest in AI alignment and reducing AGI risk. He is investigating dangerous capabilities and the emergence of misalignment in LLMs, along with self-awareness and latent reasoning. Owain previously worked on AI deception (How to Catch an AI Liar), truthfulness (TruthfulQA), and the Reversal Curse. Owain runs an independent AI Safety non-profit, based at Constellation in Berkeley. He previously worked at the University of Oxford and at Ought. He has mentored 30+ junior AI Safety researchers through MATS and other programs.

Read more
Jan Betley
Truthful AI
,
Researcher
Misalignment Science

Jan worked as a software developer for over a decade before shifting to AI safety in 2023. He is an ARENA and Astra Fellowship alumni, interested in anything related to out-of-context reasoning in LLMs.

Read more

Project selection