Daniel Kang

I have two broad areas.

Security:

I am interested in building demonstrations for hacking real-world AI deployments to show that they are not secure. The goal is to force companies to invest in alignment techniques that can solve the underlying security issues.

Verification:

Verification via TEEs or ZKPs

Stream overview

For security:

You will focus on hacking real-world AI deployments to show that they are not secure. 

For verification: TEEs or ZKPs

Mentors

Daniel Kang
University of Illinois Urbana-Champaign
,
Professor
SF Bay Area
Biosecurity
Capability and Propensity Evaluations
AI Systems Security
Technical AI Governance
Adversarial Robustness and Safeguards

Daniel is a professor of computer science at UIUC, where he studies the progress of AI, with a particular focus on dangerous capabilities of AI agents. His work includes:

  • CVE-Bench, an award winning benchmark (SafeBench award, ICML spotlight) that is used by frontier labs and governments to measure AI agents' ability to find and exploit real-world vulnerabilities.
  • Agent Benchmark Checklist, an award winning work (Berkeley AI summit, 1st place Benchmarks & Evaluations track) that highlights major issues in existing benchmarks.
  • InjecAgent, one of the first AI agent safety benchmarks, used by governments and major labs.

Read more

Mentorship style

I will meet 1-1 or as a group, depending on the interests as they relate to the projects. Slack communication outside of the 1-1.

I strongly prefer multiple short meetings over single long meetings, except at the start.

I'll help with research obstacles, including outside of meetings

Fellows we are looking for

For security:

You should have a strong security mindset, having demonstrated the willingness to be creative on this. I would like to see past demonstration of willingness to get your hands dirty and try many different systems.

For benchmarks:

As creative as possible, willingness to work on the nitty gritty, willingness to work really hard on problems other people find boring. Interests as far away from SF-related interests as possible.

Project selection

Mentor(s) will talk through project ideas with scholar

Streams

The Winter 2027 cohort offers a wide range of research streams led by experts across AI alignment, interpretability, governance, and safety. Each stream provides its own research agenda, methodology, and mentorship focus.

London
Empirical
Interpretability
London
Interpretability, Red-Teaming, Monitoring
London
Monitoring, Adversarial Robustness, Control, Model Organisms, Red-Teaming, Dangerous Capability Evals, Safeguards
New York City
Dangerous Capability Evals, Control, Strategy & Forecasting, Policy & Governance, Scalable Oversight, Agent Foundations
SF Bay Area
Empirical
Theory
Dangerous Capability Evals, Adversarial Robustness, Security, Red-Teaming, Scalable Oversight
London
Control, Scheming & Deception, Dangerous Capability Evals, Monitoring
Washington, D.C.
Policy & Governance, Strategy & Forecasting
Oxford
AI Welfare
SF Bay Area
Control, Model Organisms, Scheming & Deception, Strategy & Forecasting
SF Bay Area
Interpretability
Tübingen
Dangerous Capability Evals, Agent Foundations, Adversarial Robustness, Monitoring, Scalable Oversight, Scheming & Deception
SF Bay Area
Dangerous Capability Evals, Policy & Governance
New York City
Monitoring, Dangerous Capability Evals, Scalable Oversight, Safeguards
SF Bay Area
Strategy & Forecasting, Policy & Governance
Montreal
Agent Foundations, Dangerous Capability Evals, Monitoring, Control, Red-Teaming, Scalable Oversight
SF Bay Area
Control, Model Organisms, Red-Teaming, Scheming & Deception