AI Governance Library

Risk Reporting for Developers' Internal AI Model Use

Given the pace of AI R&D automation and the limited external visibility into how companies use their most capable models internally, regular and detailed risk reporting may be one of the few mechanisms available to ensure that the risks from internal AI use are identified and managed
Risk Reporting for Developers' Internal AI Model Use

⚡ Quick Summary

Frontier AI developers frequently deploy their most capable models internally for weeks or months prior to external release—or keep specialized unaligned variants exclusively internal. This practice introduces serious risks that standard external deployment governance frameworks fail to address. Published by the Institute for AI Policy and Strategy (IAPS), this guide proposes a harmonized reporting standard to satisfy emerging regulatory mandates under California's SB 53, New York's RAISE Act, and the EU AI Act's GPAI Code of Practice. The authors structure internal risk assessments around a safety case framework addressing two primary threat vectors: autonomous AI misbehavior and rogue insider threats. For each vector, frontier developers must evaluate three risk factors—means, motive, and opportunity—and demonstrate that internal deployment poses no meaningfully greater marginal risk compared to publicly available AI systems.

🧩 What's Covered

The report details a comprehensive structure for internal AI risk evaluation and regulatory disclosure across several key areas:

  • Distinct Internal Risks: Internal models possess privileged access to infrastructure, codebases, and weights; lack external scrutiny; and often feature heightened capabilities or helpful-only configurations without safety filters. Deep integration into automated AI R&D compounds potential harm.
  • Threat Vectors and Harmful Outcomes: The guide examines two core threat vectors—autonomous AI misbehavior (such as sandbagging, alignment faking, or covert reasoning) and insider threats (employees or contractors abusing privileged access). These vectors drive three major catastrophic outcomes: sabotage of internal AI R&D, direct catastrophic misuse (e.g., biological or cyber offenses), and model weight exfiltration.
  • The Safety Case Framework: Developers must build structured, contestable arguments evaluating means (technical capability to execute harm or evade detection), motive (inclination toward misalignment or rogue actions), and opportunity (gaps in access controls, monitoring, and containment). Safety arguments must demonstrate acceptable marginal risk relative to public baselines.
  • Reporting Scope and Cadence: A clear decision tree defines when systems are in scope (more capable, fewer safety features, or novel unmonitorable training architectures). Invalidation of a safety case requires full or targeted updates within 14 days, supported by routine quarterly backstop reports.
  • Evidence and Evaluation Standards: The guide specifies concrete metrics, including agentic benchmarks (RE-Bench, SWE-Bench Pro), real-world R&D automation indicators, honeypot evaluations, sparse autoencoder interpretability, model organisms, uplift testing against CBRN benchmarks, and granular tier-based access controls.
  • Reporting Safeguards: Recommendations require proportional reporting depth, explicit counterarguments, bounded uncertainty, explained omissions, and confidential regulatory annexes for sensitive insider threat defenses to prevent security roadmaps from leaking.

💡 Why it matters?

As automated AI research accelerates, frontier developers increasingly rely on unreleased, highly capable models with direct access to proprietary codebases and training pipelines. External deployment disclosures overlook these closed environments. By harmonizing compliance requirements across California's SB 53, New York's RAISE Act, and the EU Code of Practice into a single contestable safety case standard, this framework enables consistent cross-developer evaluation, equips regulators to monitor hidden capability leaps, and ensures organizations systematically mitigate internal sabotage, weight exfiltration, and misuse risks.

❓ What's Missing

The guide explicitly excludes outsider threats operating without insider collusion and sets aside non-deliberate capability failures, such as unintentional software bugs. Furthermore, it notes that measuring model motive remains deeply challenging due to deceptive alignment and alignment faking. While recommending specific organizational and technical metrics, the authors acknowledge that operationalizing real-world R&D automation contributions lacks standard definitions and heavily depends on regulatory capacity to evaluate submitted evidence.

👥 Best For

Evaluation and safety teams at frontier AI labs, AI governance and compliance officers, technical risk auditors, and regulatory bodies (such as the European AI Office and state regulators) responsible for overseeing frontier model development pipelines and internal deployments.

📄 Source Details

Delaney, O., Maheshwari, S., O'Brien, J., Bearman, T., & Guest, O. (2026). Risk Reporting for Developers' Internal AI Model Use. Institute for AI Policy and Strategy (IAPS). Published April 2026.

📝 Thanks to

Charles Foster, Zaheed Kara, Tom Reed, Matteo Pistillo, Kathrin Gardhouse, Girish Sastry, Leonie Koessler, Daniel Kokotajlo, Joe Kwon, Aidan Homewood, Lily Stelling, Thomas Woodside, and Michael Chen for their discussion and input.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.