WHITZARD INDEX · NUWA

Track frontier AI risk over time

A longitudinal evidence system for frontier AI risk

Whitzard Index organizes evaluation results into longitudinal, comparable, and citable evidence across models and agent systems.

01
Preserve dimensions
02
Bound conclusions
03
Cite original evidence
FRONTIER RISK OBSERVATORY
PUBLISHED EVIDENCE · NOT A COMPOSITE RANKING
Evaluation scope
  • 110 real-world web vulnerabilities
  • 8 enterprise network environments
  • 156 internal hosts
Published Last verified
GPT-5.6 Sol + Codex (0.144.1)Web exploitation success
34.53%Latest verified configuration
GPT-5.6 Sol + Codex (0.144.1)Post-exploitation success
74.10%Latest verified configuration
Finding

The same agent configuration shows materially different capability across web exploitation and enterprise post-exploitation tasks.

Interpretation boundary

Results apply only to the listed model, Codex version, task set, and submission date; they are not a general probability of cyberattack success.

Original source: AgentCyberRange

About the index

Multidimensional risk evidence, not a single safety score

Whitzard Index preserves the meaning, scope, and uncertainty of each risk dimension. Published evidence can be followed over time, while results produced by different methods remain visibly separated rather than being collapsed into a misleading overall ranking.

Positioning

Principles for credible risk evidence

Multidimensional evidence

Risk dimensions retain their individual meaning instead of being collapsed into an unjustified total score.

Separated comparisons

Raw-system results and protected-system results are reported separately with clear labels and configuration context.

Reproducible records

Each evaluation carries a method version, system configuration, date, and evidence source.

Bounded conclusions

Findings remain scoped to the systems, environments, permissions, and metrics that were actually tested.

Methodology

Evidence record design

The information each Whitzard Index record must preserve for results to remain interpretable and citable.

Evaluation subject

The exact model or agent system, version, and evaluation date.

Runtime configuration

Agent scaffold, tool permissions, resource budgets, and other conditions that can change the result.

Risk dimension

The specific behavior or capability measured, with its unit and interpretation direction.

Method version

The method used to produce the result and the changes that affect comparability over time.

Evidence source

The original paper, project, or dataset that supports the published observation.

Uncertainty

Known limits, missing coverage, and the boundary beyond which a conclusion should not be generalized.

Any published result must identify its method version and preserve uncertainty rather than imply broader coverage.

COLLABORATION

Whitzard Index collaboration

For inquiries about the Whitzard Index — methodology, data contribution, or research collaboration — contact NUWA.