Alex Cloud

Anthropic

—

Member of Technical Staff (Alignment Science)

Links

Focus

Misalignment Science, Alignment Training Methods, Interpretability

Alex is a researcher at Anthropic. He is interested in developing principled methods to induce safety-relevant structure in models. Examples include gradient routing to localize learning updates in models and distillation for robust unlearning.

​

Previously, Alex conducted applied research in reinforcement learning at Riot Games AI and Amazon. He earned a PhD in Statistics from North Carolina State University, where he was advised by Eric Laber.