I am broadly interested in research directions scholars are excited about, that can advance the quality of our AI Safety tools, and our confidence in them. Three particular areas of research that seem promising to me are:
London
Arthur Conmy is a Member of Technical Staff at Anthropic. His interests are in automating interpretability, finding circuits and making model internals techniques useful for AI Safety, particularly with Sparse Autoencoders. Previously, he worked at Google DeepMind and Redwood Research (and did the MATS Program!).
Executing fast on projects is highly important. But also having a good sense of which next steps are correct is also valuable, though I enjoy being pretty involved in projects, so it's somewhat easier for me to steer projects than it is for me to teach you how to execute fast from scratch. It helps to be motivated to make interpretability useful, and use it for AI Safety, too.
I will also be interviewing folks doing Neel Nanda's MATS research sprint who Neel doesn't get to work with.
Mentor(s) will talk through project ideas with scholar.