Arthur Conmy

Anthropic

—

Research Engineer

Links

Focus

AI Control and Monitoring, Alignment Training Methods, Interpretability

Arthur Conmy is a Member of Technical Staff at Anthropic. His interests are in automating interpretability, finding circuits and making model internals techniques useful for AI Safety, particularly with Sparse Autoencoders. Previously, he worked at Google DeepMind and Redwood Research (and did the MATS Program!).