The Alignment Research Center is a small non-profit research group based in Berkeley, California, that is working on a systematic and theoretically grounded approach to mechanistically explaining neural network behavior. We are interested in scholars with a strong math background and mathematical maturity. If you'd be excited to work on the research direction described in this blog post – then we'd encourage you to apply!
ARC will be supervising projects that fit into our technical research agenda, which is outlined here. Such projects could be:
SF Bay Area
Eric Neyman is a researcher at the Alignment Research Center (ARC), which is working on a systematic and theoretically grounded approach to mechanistic interpretability. Before joining ARC, he was a PhD student at Columbia University, where he researched algorithmic Bayesian epistemology.
Wilson Wu is a researcher at the Alignment Research Center (ARC), which is working on a systematic and theoretically grounded approach to mechanistic interpretability. He has previously worked on alternate approaches to interpretability including compact proofs and applications of singular learning theory.
Victor Lecomte is a researcher at the Alignment Research Center (ARC), which is working on a systematic and theoretically grounded approach to mechanistic interpretability. He holds a PhD from Stanford University, where he did research in computational complexity and other areas of theoretical computer science before pivoting to AI safety research.
George Robinson is an independent researcher formerly at the Alignment Research Center (ARC), working on a systematic and theoretically grounded approach to mechanistic interpretability. He is now looking to lead a research effort in London supporting this agenda. Previously, he was a PhD student at Oxford University specialising in Algebraic Number Theory. He lives in London, and is a member of the London Initiative for Safe AI (LISA).
Jacob Hilton is a researcher at the Alignment Research Center (ARC), a nonprofit working on the theoretical foundations of mechanistic interpretability. He previously worked at OpenAI on reinforcement learning from human feedback, scaling laws and interpretability. His background is in pure mathematics, and he holds a PhD in set theory from the University of Leeds, UK.
Mike Winer is a researcher at the Alignment Research Center (ARC), where he studies how mechanistic estimates can beat black-box techniques in toy setups. His background is in statistical physics, where he studies how many objects obeying simple rules can exhibit complex behaviors like magnetism, glassiness, or scoring 87% on GPQA.
Essential:
Preferred:
Each scholar will be paired with the mentor that best suits their skills and interests. The mentor will discuss potential projects with the scholar, and they will decide what project makes the most sense, based on ARC's research goals and the scholar's preferences.
Most scholars will work on multiple projects over the course of their time at ARC, and some scholars will work with multiple mentors.