⚡ Quick Summary
Conducted by the Forecasting Research Institute, this pilot study applies structured forecasting with 21 participants—13 superforecasters and 8 cybersecurity experts—to evaluate near-term catastrophic AI cyber risks in 2026. The research investigates two primary threat vectors: large-scale data-damaging worm attacks (modeled on WannaCry and NotPetya) and cyberattacks against the U.S. electrical grid causing blackouts with at least $10 billion to $100 billion in damages. Baseline risks were evaluated as low but non-negligible (5–8% for $10B+ worms; 1% for $10B+ grid blackouts). Crucially, the introduction of AI-enabled elite exploit development by moderately-skilled actors (TA2) increased worm attack risk estimates by 3 to 3.5 times, raising expert probability estimates to 41% and superforecasters to 15%. While proprietary model access controls and anti-jailbreaking mechanisms more than halve these risks, grid attacks remain heavily constrained by operational complexity, physical infrastructure barriers, and state-level geopolitical deterrence.
🧩 What's Covered
The report evaluates structured elicitation across threat actor tiers (TA1 hobbyists to TA5 top nation-states) and specific attack pathways:
- Data-Damaging Worm Scenarios: Assessed baseline probabilities of a ≥$10B damage attack at 8% (experts) and 5% (superforecasters), with expected annual damages of $10–$15 billion. The primary operational bottleneck identified is the creation of zero-click, high-privilege remote code execution "elite exploits."
- AI Exploit Uplift (Capability 1): When open-weight frontier models enable 25% of moderately skilled actors (TA2) to discover vulnerabilities and write elite exploits within three months, median expert risk rises to 41% (from 8%) and superforecasters to 15% (from 5%), expanding expected annual damages to $33–$67 billion.
- Mitigation Efficacy (Policies P1 & P2): Proprietary API-based distribution with red-teaming, universal jailbreak patching within two weeks, bug bounties ($15k), and RAND Security Level 2 controls (P1) more than halved worm probabilities. Providing temporary early access exclusively to defenders (P2) was deemed less effective by experts due to short remediation windows and the massive defense surface.
- U.S. Electrical Grid Threats: Baseline risks for ≥$10B outages were 1% ($0.2B–$1B expected damages) and 0.1% for ≥$100B outages. AI uplift on operational technology (OT/ICS) capture-the-flag competitions (Capability 3) only modestly increased risk (1.5–2%), because grid attacks require physical reconnaissance, bespoke SCADA/ICS knowledge, and human coordination. However, an AI-assisted $100M "warning shot" (Capability 4) increased expert risk to 15%.
- Benchmark Validity & Real-World Translation: Evaluated Cybench (>90% solve rate) as an incomplete proxy for real-world risk, failing to measure zero-day discovery, persistence, targeting, and evasion. Follow-up survey analysis of Anthropic's reported Claude-assisted espionage incident revealed that automated multi-stage attacks still rely heavily on known vulnerabilities rather than elite exploit synthesis.
💡 Why it matters?
For AI governance and risk leaders, this study provides rigorous empirical calibration separating speculative cyber fears from actual threat mechanics. It shows that AI's greatest near-term cyber risk does not stem from sophisticated nation-states (who are constrained by retaliation red lines), but from lowering technical barriers for abundant, unconstrained mid-tier actors (TA2) to generate zero-click exploits. It also offers quantitative backing for frontier AI safety policies—confirming that API guardrails, structured access, and model weight security significantly suppress risk escalation.
❓ What's Missing
As an exploratory pilot, the study relies on a small convenience sample (21 participants, narrowing in follow-ups), which makes aggregate quantiles fragile and subject to selection bias. Additionally, abstract conditional forecasting questions lack retrospective resolvability scoring. The study does not evaluate non-kinetic cyber harms like systemic data exfiltration or autonomous financial fraud, nor does it deeply model defensive AI acceleration beyond a four-month exclusive defender window.
👥 Best For
AI safety researchers, frontier model deployment and security teams, critical infrastructure CISOs, threat intelligence analysts, and national security policymakers seeking structured empirical assessments of dual-use AI capabilities in offensive cyber operations.
📄 Source Details
- Title: Forecasting AI Cyber Risks and Capabilities: Results of a 2025 Pilot Study
- Authors: Rebecca Ceppas de Castro, Bridget Williams, John Halstead, Matthew van der Merwe, Josh Connor
- Organizations: Forecasting Research Institute (FRI Report #7), University of Oxford, Centre for the Governance of AI (GovAI)
- Publication Date: July 23, 2026
- Methodology: Delphi-style structured elicitation and follow-up surveys with 13 superforecasters and 8 domain experts
📝 Thanks to
Reviewed by Kuba Szarmach for the AI Governance Library.