
Anthropic
—
Research Engineer
Links
Focus
AI Control and Monitoring, Alignment Training Methods, Interpretability
Stream
Arthur Conmy
Arthur Conmy is a Member of Technical Staff at Anthropic. His interests are in automating interpretability, finding circuits and making model internals techniques useful for AI Safety, particularly with Sparse Autoencoders. Previously, he worked at Google DeepMind and Redwood Research (and did the MATS Program!).