I mostly interested in AI control and scalable oversight. I'm excited to work with scholars interested in empirical projects building and evaluating control measures and oversight techniques for LLM agents, especially those based on chain of thought monitoring. I'm also interested in the science of chain of thought monitorability, misalignment and control. An ideal project ends with a paper submitted to NeurIPS/ICML/ICLR.
Some example projects:
I'll meet with mentees once a week and will be available on Slack daily. By default, I'll try to pair mentees to work on a project together.
I'll meet with mentees once a week and will be available on Slack daily.
An ideal mentee has a strong AI research and/or software engineering background. A mentee can be a PhD student and they can work on a paper that will be part of their thesis.
By default, I'll try to pair mentees to work on a project together.
I'll talk through project ideas with scholar