MATS Fellow:
Jennifer Nzambi, Ares Panos
Authors:
Jennifer Za, Aristeidis Panos, Jan Cuhel
Citations
Abstract:
Large language models (LLMs) are increasingly deployed as autonomous agents in settings involving cooperation, competition, and coordination, yet current behavioural evaluations provide limited guidance for anticipating risks in deployment. We present a large-scale study of strategic decision-making across seven frontier models, analysing over 200,000 decisions in game-theoretic scenarios. Using controlled experiments, we found that apparent self-recognition effects operate through inferred policy correlation rather than identity; a correlated stranger elicits cooperation equivalent to a correlated self. We further observe substantial heterogeneity across model families, including opposite responses to identical ``rationality'' instructions, which one might use to steer agent behaviour, and marked differences in forgiveness and exploitation dynamics in iterated interactions. Finally, we introduce a lightweight prediction method that requires only 5-10 calibration scenarios and achieves R^2 = 0.51 on average (up to R^2 = 0.70) when forecasting held-out model behaviour. These results demonstrate that systematic behavioural evaluation of LLMs can support pre-deployment risk assessment and shed light on AI agent decision-making in strategic situations.
Synthetic Persona Pretraining: Alignment from Token Zero
Authors:
Julian Minder
Date:
August 13, 2026
Citations:
The MATS Program is an independent research and educational initiative connecting emerging researchers with mentors in AI alignment, governance, and security.
Each MATS cohort runs for 12 weeks in Berkeley, California, followed by an optional 6–12 month extension in London for selected scholars.