Counterfactual Debugging the World Model Transfer Gap

MATS Fellow:

Mingxuan Li

Authors:

Mingxuan Li, Kai-Zhan Lee, Michael Dennis, Elias Bareinboim

Citations

Citations

Abstract:

Policies achieving strong performance in simulators or learned world models can fail when deployed into the real environment. While there must be environmental differences that account for these failures, not all differences are equally responsible. A benign error of irrelevant visual details may co-exist with a failure to predict a single critical transition causing an inevitable catastrophic failure. In this paper, we introduce counterfactual debugging as an approach for world model transfer gap attribution to identify the time steps whose transition or reward errors are causally responsible for a performance degradation exhibited in a real trajectory, rather than only visually or statistically different. We develop a scalable algorithm for computing these attributions that exploits the sparsity of causal errors by recursively divide-and-conquer, significantly reducing computational costs and achieving an exponential speedup. The resulting ranked attribution report can better explain the performance gap in the real trajectory. Experiments across several distinct environments with a variety of injected world-model failures (observation corruption, reward misprediction, and physics violations) demonstrate that counterfactual debugging correctly identifies the errors that are responsible for performance gap, providing theoretically grounded, actionable insights for model improvement.

Recent research

Counterfactual Debugging the World Model Transfer Gap

Authors:

Mingxuan Li

Date:

December 8, 2026

Citations:

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Authors:

Joachim Schaeffer, Alexander Panfilov

Date:

October 5, 2026

Citations:

Frequently asked questions

What is the MATS Program?
How long does the program last?