Alex Turner

Independent

Independent researcher

Links

Focus

Alignment Training Methods, Interpretability, Agent Foundations

Alex is currently working on training invariants into model behavior. In the past, he formulated and proved the power-seeking theorems, co-formulated the shard theory of human value formation, and proposed the Attainable Utility Preservation approach to penalizing negative side effects.

Highlighted outputs from past streams: