AI Governance Library

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management

SaferAI paper proposing a four-part frontier AI risk management framework — risk identification, analysis and evaluation, treatment and governance — built around an explicit risk tolerance operationalized into KRI and KCI thresholds linked by if-then rules.
Cover of A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk…

⚡ Quick Summary

Authored by researchers at SaferAI and circulated as arXiv preprint 2502.06656v3, this paper presents a four-part risk management framework for the development of frontier AI, assembled by combining established practices from safety-critical industries with emerging AI-specific methods. Its purpose is to close the gap between the safety frameworks AI developers have begun to publish and the systematic rigor found in sectors such as aviation and nuclear power.

The framework comprises risk identification, risk analysis and evaluation, risk treatment, and risk governance. Its central mechanism is to define a risk tolerance, operationalize it into measurable Key Risk Indicators (KRIs) and Key Control Indicators (KCIs) linked by "if-then" thresholds, and implement mitigations whenever KRI thresholds are crossed. The paper explains how each component fits the AI life-cycle, arguing that most work—risk modeling, threshold setting, and predicting required mitigations with scaling laws—can be done before the final training run, leaving KRI measurement and open-ended red teaming for training and pre-deployment.

The conclusion states that the explicit setting of a quantitative risk tolerance is the element most missing from current practice and identifies risk modeling and quantitative assessment as the main methodological gaps.

🧩 What’s Covered

The paper proceeds from a review of current practice through the four framework components to implementation guidance, closing with limitations and reference material.

  • Background and motivation: reviews developer safety frameworks such as Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework, and the Frontier Safety Commitments from the 2024 Seoul AI Summit, arguing these lack a defined risk tolerance, quantitative assessment and systematic risk identification.
  • Risk identification: covers classification of known risks from taxonomies and literature (Weidinger et al., the AI Risk Repository), open-ended red teaming conducted internally and by third parties to surface unknown risks, and risk modeling using Probabilistic Risk Assessment with event and fault trees, Fishbone diagrams for triage, and documented scenarios.
  • Risk analysis and evaluation: explains setting a risk tolerance, quantitative (probability times severity per unit of time) or as qualitative scenarios with quantitative probabilities, then operationalizing it into KRI and KCI thresholds joined by an if-then relationship; examples include containment, deployment and assurance-process KCIs and a fictional Cybench threshold illustration.
  • Risk treatment: sets out containment measures, deployment measures and assurance processes, continuous monitoring of KRIs and KCIs, evaluation protocols producing upper-bound capability estimates, external vetting of protocols and sharing of results with stakeholders.
  • Risk governance: describes six categories—decision-making, advisory and challenge, culture, oversight, audit and transparency—with roles including the risk owner, Chief Risk Officer, Enterprise Risk Management, board audit or risk committees, and internal and external auditors.
  • Implementation: proposes a continuously updated risk register covering risk owner, risk level, KRIs, mitigation status and KCIs, KRI/KCI mapping and an action plan, and places the framework's components across the planning, training and post-deployment phases.
  • Conclusion and reference material: restates the four components and the framework's limitations, and adds an abbreviations list, a glossary (assurance processes, capabilities thresholds, risk tolerance, scaling laws) and the reference list.

💡 Why it matters?

For AI developers, the framework supplies machinery that higher-level commitments lack: a documented risk tolerance, measurable KRI and KCI thresholds, and an if-then rule that ties mitigation to capability. It gives risk, security and audit functions a structure for deciding who owns a risk, who challenges decisions, and who independently tests the framework, plus a life-cycle plan that front-loads the work before the final training run. The paper connects its components explicitly to the NIST AI RMF, Raz and Hillson's steps, and standards such as ISO/IEC 42001 and ISO/IEC 23894.

❓ What’s Missing

The paper is candid about its gaps. It states that frontier AI still lacks a detailed understanding of how harms materialize, and that current quantitative risk assessment methods are insufficient to demonstrate rigorously that KCI thresholds keep risk below tolerance; a footnote notes that establishing quantitative links between capabilities and risks remains challenging. Assurance processes able to provide high-safety guarantees for large language models have not been demonstrated. The framework argues regulators should ideally set risk tolerance but assumes developers act voluntarily meanwhile, and it gives no cost, staffing or verification detail for the governance structure.

👥 Best For

Best suited to safety, risk and governance teams at frontier AI developers drafting or revising internal safety frameworks; to boards and audit or risk committee members who need to know what oversight of AI risk involves; and to auditors, policy analysts and regulators comparing developer commitments with established risk management practice.

📄 Source Details

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management, by Siméon Campos, Henry Papadatos, Fabien Roger, Chloé Touzet, Otter Quarks and Malcolm Murray of SaferAI. Circulated as arXiv preprint arXiv:2502.06656v3 [cs.AI], dated 19 February 2025; 20 pages; English; the text extraction covered all 20 pages. It includes an abbreviations list, a glossary and a reference list. No URL for the document itself is printed.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.