Neev Parikh

This stream focuses on empirical projects that improve our ability to evaluate model capabilities or enable us to understand or evaluate model monitorability. An ideal project culminates in a research output (conference/Arxiv paper or research blogpost with artifacts).

Stream overview

Potential projects examples include:

  • Build an evaluation measuring how good models are at predicting which research directions/approaches are better than others, and how they compare to human predictors
  • Try to compute a "Judge" horizon, which aims to understand the trend of models accurately verifying/predicting the score or result of an evaluation transcript. The goal would be to determine if we can plot the discriminator-generator gap and understand if the kinds of tasks that model N can verify are the kinds of tasks model N+1 gains the most improvement on.
  • Models could likely use additional inference time compute more effectively for a number of current benchmarks. A potential approach is trying to quantify and reduce this gap by applying existing techniques that target specific weaknesses or experiment with new techniques.

These project ideas are less well-scoped by me, but I'd be interested and excited if fellows have clear ideas of what to work on here:

  • Build better red-team attack policies for monitorability evaluations.
  • Models are increasingly trained with synthetically generated data. Does this mean that models will be able to instinctively "underperform" (or sandbag) by behaving like the less-capable model that generated their training data without needing to reason about this?
  • Understand if emergent misalignment behaves similarly with RL training versus fine-tuning (e.g. do you see more reward hacking if you RL models in CTF or reverse engineering environments?).
  • Try to understand the relationship between compute and time-horizon trends. METR's time horizon measurements use release date on the x-axis instead of effective compute. Could we try and measure how effective compute relates with time-horizon?
  • Try and understand how various recently proposed neuralese/latent reasoning architectures compare to the current open-source default reasoning paradigm and estimate how the alignment tax for visible-chain-of-thought has changed over time. This might look like compiling the presented results and analyzing trends but also might involve collecting new data by benchmarking the relevant models.
Mentorship style:

Light-touch (.5 hours of weekly 1:1s)

Location during program:

SF Bay Area

London location preference:

No preference for this location

Berkeley location preference:

No preference for this location

Mentors

Neev Parikh
METR
,
Member of Technical Staff
Misalignment Science
AI Control and Monitoring
Capability and Propensity Evaluations

Neev likes to make computers do interesting things, deeply understand concepts and build interesting, useful tools. He is currently thinking about AI alignment, control, and evaluations, and work with frontier models at METR.

Recent work done involves MALTtraining models to fool monitors in QA settings and RE-Bench. Neev previously worked at Stripe and CSM, and did a concurrent BSc/MSc in Computer Science at Brown.

Read more

Fellows we are looking for

An ideal candidate has a strong AI research background; software engineering is a plus. It's important that they are self-motivated and can make weekly progress with little intervention. If you are interested in working on non-concretely scoped projects, I would expect candidates to have the ability to write well-scoped project proposals, with realistic planned milestones and deliverables. Evidence of successful projects here would be very helpful in evaluating this.

A candidate that’s a PhD student can work on a paper that will be part of their thesis.

Project selection

I will talk through project ideas with the scholar