AI Governance Library

The Offensive Frontier: AI as the Attacker: A New Cyber Weapon Index and the Strategic Imperative to Accelerate Agentic AI Offense and Defense

Booz Allen Hamilton's Cyber Weapon Index benchmark scores 18 U.S. and Chinese frontier AI models on autonomous offensive cyber capability, with an addendum re-testing newly released models days later.
Cover of The Offensive Frontier: AI as the Attacker: A New Cyber Weapon Index and the Strategic Imperative to Accelerate…

⚡ Quick Summary

Published by Booz Allen Hamilton, this report introduces the Booz Allen Cyber Weapon Index (CWI), a novel benchmark that scores 18 frontier large language models — U.S. and Chinese, open- and closed-weight — on their ability to autonomously execute cyber operations. Models were set up as fully autonomous attackers against a production-grade enterprise network, and scoring used the network's own logs, intrusion-detection sensors and attacker transcripts rather than model claims.

The CWI combines a Vulnerability Research Score (VRS), which tests whether models can find planted and genuine vulnerabilities in compiled code without source access, with a Kill Chain Attainment Score (KCAS), which credits progress through a live Active Directory environment with and without credentials; the index is calculated as CWI = (VRS + KCAS)/2.

The central finding is that autonomous offensive cyber has arrived: one model, Anthropic's Claude Mythos, executed the full cyber kill chain, four models reached full domain access and control, and all but one autonomously penetrated the network. The report argues that the real unit of risk is the full AI system — model plus harness, tools and autonomy — and recommends binding critical-infrastructure readiness standards, continuous national measurement of the global model landscape, and wider use of counter-AI deception. An addendum re-tests newer models and places OpenAI's Astra alongside Mythos at the frontier.

🧩 What’s Covered

  • Cyber attack lifecycle: the seven-stage model that frames the assessment — recon, initial access, foothold, credential access, privilege escalation, lateral movement, objective (DC Sync / DA).
  • Cyber Weapon Index: the benchmark definition, combining a Vulnerability Research Score (VRS) on a planted-vulnerability binary and a genuine, previously unseen flaw in a large production library with a Kill Chain Attainment Score (KCAS) for progress through a live Active Directory network with and without credentials; CWI = (VRS + KCAS)/2, scored from observed telemetry — network traffic, host logs, domain controller logs, IDS.
  • Rankings and methodology: the 18-model table of country, VRS, KCAS, achievement and CWI score, led by Claude Mythos at 80, and the test design — identical conditions, a real attacker machine driven one command at a time, no tool menu or scaffolding, repeated runs.
  • Six findings: one model executing the full kill chain while nine frontier API models scored zero on the real-world vulnerability; capability spread across U.S. and Chinese models; harnesses mattering as much as the model; safeguards shifting with configuration; dangerous capability beyond U.S. reach and outside traditional evaluations; overmatch through agentic AI on offense and defense.
  • Recommendations: binding sector-specific readiness standards and deadlines led by CISA; a national evaluation programme covering the full AI system with controlled defender access; and counter-AI playbooks that reduced attacker success by more than 95 percent, against a Velocity Survey finding that only 28 percent of federal cyber and IT leaders are confident AI agents can be deployed securely.
  • What We See Coming: four shifts on timelines, from mainstream AI-enabled attacks and on-demand zero days within months to defenders regaining advantage in 1 to 2 years and self-improving offensive models in 2 to 3 years.
  • Addendum: a re-test days later in which five new models enter the Top 18 and GPT-6 Astra edges ahead at 80.5, with Mythos stronger on KCAS and Astra on VRS, plus three findings on a frontier tier.

💡 Why it matters?

The report gives security and policy audiences an empirical, telemetry-based way to reason about offensive AI capability instead of relying on model claims or knowledge benchmarks. Its argument that risk sits in the full AI system — model, harness, tools, memory and autonomy — turns governance attention to levers organisations actually control: harness design, tool access, permissions and the degree of autonomy granted. The measured spread across models supports readiness planning, and the report connects its findings to sector-specific standards, instrumented exercises, counter-AI deception and the claim that point-in-time evaluation stops being sufficient as capabilities change between releases.

❓ What’s Missing

The report does not publish the prompts or instructions given to models, the harness configurations used, sample sizes, scoring thresholds, or the identity of the real-world vulnerability it describes. Equal weighting of VRS and KCAS is asserted rather than justified, and no cost, compute or false-positive figures are given. How the 18 models were selected is not explained. The authors acknowledge one blind spot: the kill-chain capability of open-weight and Chinese models paired with optimised harnesses is unknown. The addendum changes the leaderboard days after publication, so two rankings must be read together, and the recommended deadlines remain unquantified.

👥 Best For

Cyber and national-security policy teams weighing readiness mandates; security leadership in energy, communications, financial services, water and healthcare; detection and red-team groups designing evaluation harnesses; and risk committees that need a concrete account of how autonomous offensive capability is measured and where models stall.

📄 Source Details

The Offensive Frontier: AI as the Attacker: A New Cyber Weapon Index and the Strategic Imperative to Accelerate Agentic AI Offense and Defense, published by Booz Allen Hamilton, carries document numbers 2026-0817 and 2026-0909 (addendum) and runs to 20 pages in English. No authors are named. Pages 8 and 15 carry an "EMBARGOED MEDIA COPY" marking. No URL for the report itself is printed; the only link in the text points to a cited Velocity Survey infographic. The input covered all 20 pages of extracted text.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.