AI Village

This stream works on the AI Village, where more than twenty frontier agents act in the real world over thousands of hours. Projects either mine the existing record of agent reasoning, messages and actions for new insight, or run new experiments in the Village to test hypotheses about multi-agent dynamics and real-world evaluation.

Stream overview

Anything related to the AI Village. I'd be especially excited for research projects where people either analyze our existing data for new insights or run their own version for the Village to test specific hypotheses. The experiments could be around multi-agent dynamics or real-world evaluation.

Specific examples:

  • Test a better AI monitoring system by using our 1000s of hours of CoTs, messages, and actions of over 20 different agents acting in the real world
  • Look for traces of steganography or deception in the same existing data
  • Design and run a real-world eval where multiple frontier models are shown to perform better in unison than alone through self-organization
  • Design and run a real-world eval that shows agent-to-agent jailbreaking or misalignment (e.g., where a more capable and aligned agent is "corrupted" by a less capable and less aligned one in a surprising way)

These ideas are illustrative of what's possible, but we'd be excited to hear new ideas we haven't thought of!

Notably: Mentees can run their experiments using our compute budget and setup if we approve their project.

Mentors

Shoshannah Tekofsky
Sage
,
Member of Technical Staff
No items found.

Shoshannah Tekofsky is a member of technical staff at Sage, the nonprofit that runs AI Digest and the AI Village. Tekofsky works on the AI Village, in which several frontier language models pursue their own goals continuously on their own computers.

Read more

Mentorship style

Fellows we are looking for

Project selection

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Control, Scheming & Deception, Dangerous Capability Evals, Model Organisms, Monitoring
SF Bay Area
Security, Dangerous Capability Evals
SF Bay Area
Dangerous Capability Evals
Boston
Adversarial Robustness, Policy & Governance, Red-Teaming, Safeguards
New York City
Control, Scalable Oversight, Red-Teaming, Model Organisms, Monitoring
SF Bay Area
Policy & Governance
SF Bay Area
Control, Monitoring, Dangerous Capability Evals
SF Bay Area
Security, Compute Infrastructure
London
Interpretability
London
Scheming & Deception, Dangerous Capability Evals, Control, Red-Teaming
SF Bay Area
Dangerous Capability Evals, Red-Teaming, Model Organisms, Control, Monitoring
Toronto
Interpretability
London
Control, Monitoring, Safeguards, Dangerous Capability Evals, Scheming & Deception
Chicago
Biorisk, Security, Safeguards
SF Bay Area
Interpretability, Agent Foundations
London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight