MATS Fellow:
Seoirse Murray, Allison Qi, Timothy Qian
Authors:
Seoirse Murray, Allison Qi, Timothy Qian, John Schulman, Collin Burns, Sara Price
Citations
Abstract:
LLM post-training involves many diverse datasets, each targeting a specific behavior. But these datasets encode incidental patterns alongside intended ones: correlations between formatting and content, narrow phrasings across diverse problems, and implicit associations arising from the discrete data curation process. These patterns are often invisible to developers yet salient to models, producing behaviors that surprise their creators, such as rejecting true facts presented in a particular question format. We call this chunky post-training: the model learns spurious correlations as a result of distinct chunks of post-training data. We introduce SURF, a black-box pipeline which surfaces these unintended behaviors at run time, and TURF, a tool that traces these failures back to specific post-training data. Applying these tools to frontier models (Claude 4.5, GPT-5.1, Grok 4.1, Gemini 3) and open models (Tülu 3), we show that chunky post-training produces miscalibrated behaviors, which often result from imbalanced or underspecified chunks of post-training data.
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Authors:
Agatha Duzan
Date:
August 5, 2026
Citations:
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Authors:
Anita Srinivasan
Date:
July 13, 2026
Citations:
The MATS Program is an independent research and educational initiative connecting emerging researchers with mentors in AI alignment, governance, and security.
Each MATS cohort runs for 12 weeks in Berkeley, California, followed by an optional 6–12 month extension in London for selected scholars.