You Are What You Read: Misalignment via In-Context Persona Induction

MATS Fellow:

Kyuhee Kim, Benjamin Berczi

Authors:

Kyuhee Kim, Benjamin Berczi, Cozmin Ududec

Citations

Citations

Abstract:

Broad misalignment has been produced by finetuning on narrow data, harmful or benign, and in context only by demonstrations of the undesirable behaviour itself. We show that benign data suffices in context, with no finetuning and no demonstration of harmful behaviour in the prompt. Biographical facts that converge on a single figure, placed in a model's context as ordinary conversational turns, lead it to answer as that figure on questions the facts never touch. We call this persona induction. Across nine personas and thirteen models, identity adoption rises sigmoidally with the number of facts and crosses 50% within 3 to 10 of them. Misalignment then tracks which figure is described. Harmless personas reach full adoption with near-zero misalignment, while harmful ones voice their characteristic views on unrelated questions, at rates up to 80%. A formatting instruction can gate when the persona activates. Because each fact is individually benign, accumulated biographical context is flagged by content filters on 3% of inputs against 24-33% for an equivalent direct instruction.

Recent research

You Are What You Read: Misalignment via In-Context Persona Induction

Authors:

Kyuhee Kim, Benjamin Berczi

Date:

September 6, 2026

Citations:

Non-Great-Power Conflict and AI Risk

Authors:

Kristina Kempkey

Date:

August 26, 2026

Citations:

Frequently asked questions

What is the MATS Program?
How long does the program last?