Define frontier risk
Study autonomy, deception, scheming, loss of control, and emerging agent behavior that becomes possible as capabilities grow.
NUWA · FRONTIER AI SAFETY LAB
NUWA Frontier AI Safety Lab advances frontier AI risk research and governance through reproducible evidence, evaluation methods, and public goods.
Research mandate
Frontier AI risk research and governance
Research mission
We study how increasingly autonomous AI systems create new risks — and how to measure, govern, and control those risks through reproducible evidence and open methods.
Study autonomy, deception, scheming, loss of control, and emerging agent behavior that becomes possible as capabilities grow.
Develop executable environments, benchmarks, methodologies, and public research records that make frontier risk observable and comparable.
Turn risk understanding into lightweight models for reasoning, intent, trust, and runtime defense.
Feed evidence into AgentGuard and learn from deployment feedback through a clear research-to-product boundary.
Featured research
Selected research programs defining frontier risk and practical control.
Systematically measures privacy leakage risks in general-purpose language models.
Source ↗Presents AgentDoS, a lifecycle-aware fuzzing framework for detecting resource-abuse DoS vulnerabilities in LLM-based agents.
Source ↗Extends self-replication evaluation across 32 AI systems and reports autonomous replication, self-exfiltration, adaptation, and shutdown-survival behaviors.
Source ↗RESEARCH DIRECTIONS
Study risks from increasingly capable and autonomous AI systems and develop verifiable approaches to control.
Autonomy · Deception · Scheming · Loss of control02Study systemic safety questions in reasoning, action, tool use, model behavior, and human–AI trust.
Agent behavior · Model safety · Tool use · Trustworthy interaction03Build executable, reproducible evaluations and study how evidence should inform risk assessment and governance decisions.
Capability evaluation · Risk assessment · Evaluation environments · Governance methodsWhitzard Index
The Whitzard Index is NUWA's project for continuously tracking and comparing frontier AI risk across models and systems. It is not a safety leaderboard, and it does not claim which product is best; it is a research data program for understanding how risk evolves.
Learn about the Whitzard Index →Research infrastructure
Public evaluation environments and benchmarks that make agent cybersecurity capability observable.
Evaluates frontier AI systems' autonomous cyber-attack capabilities in complex enterprise environments.
Official site ↗Synthesizes executable test environments that combine deterministic code state with narrative dynamics for frontier AI risk evaluation.
Official site ↗Complete research
Browse the complete publication record across frontier AI risk, agent safety, systems security, cybersecurity, and privacy — 86 research works with full filtering and search.
Browse all research →COLLABORATION
Connect with us on frontier-risk evaluation, agent safety, AI control, and open technical evidence.