Jacob Merizian

This stream will focus on AI control and evaluations of dangerous capabilities, propensities, and alignment.

Stream overview

I would be interested in advising work in control and dangerous capability/propensity/alignment evals, as I feel I have the best chance of being a good mentor for these sorts of projects. Currently, I work on propensity evaluations, evaluation awareness, and sandbagging. I would be happy to suggest concrete project ideas and help with brainstorming topic choices, or help guide an existing project.

Location during program:

London

Mentors

Jacob Merizian
UK AISI
,
Research Scientist, Workstream Lead
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations

I work at the UK AI Security Institute. In the past, I’ve done research in high-performance computing, language model pretraining, interpretability, and hardware enabled governance.

Read more

Fellows we are looking for

I'm open to a wider variety of skillsets, but these would be a big plus:

  • some relevant technical background in running basic finetuning/inference/interp experiments on a multi-gpu cluster
  • some prior level of interest in any of the research categories I've listed
  • ability to work independently once there is a clear enough goal (though I'm happy to be the one supplying the goal if that is the bottleneck)

Project selection

I would be happy to suggest concrete project ideas and help with brainstorming topic choices, or help guide an existing project that the scholar is interested in. My preference is that the scholar picks a category that overlaps with an area I actively work on so that I can give effective high-level advice.