I would be interested in advising work in control and dangerous capability/propensity/alignment evals, as I feel I have the best chance of being a good mentor for these sorts of projects. Currently, I work on propensity evaluations, evaluation awareness, and sandbagging. I would be happy to suggest concrete project ideas and help with brainstorming topic choices, or help guide an existing project.
London
I work at the UK AI Security Institute. In the past, I’ve done research in high-performance computing, language model pretraining, interpretability, and hardware enabled governance.
I'm open to a wider variety of skillsets, but these would be a big plus:
I would be happy to suggest concrete project ideas and help with brainstorming topic choices, or help guide an existing project that the scholar is interested in. My preference is that the scholar picks a category that overlaps with an area I actively work on so that I can give effective high-level advice.