AI Governance Library

Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems

An arXiv preprint from the AI Standards Lab presenting a catalog of risk sources and risk management measures for general-purpose AI systems across development, evaluation, deployment, cybersecurity and impact stages.
Cover of Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems

⚡ Quick Summary

Published by the AI Standards Lab as arXiv preprint arXiv:2410.23472v2 dated 15 November 2024, this paper by Rokas Gipiškis, Ayrton San Joaquin, Ze Shen Chin, Adrian Regenfuß, Ariel Gil and Koen Holtman compiles an extensive catalog of risk sources and risk management measures for general-purpose AI (GPAI) systems. It asks what short- and long-term risk sources arise from the development and deployment of GPAI models or systems, and what established and experimental methods are available to manage systemic risks at various points in the supply chain.

The catalog spans Sections 4 to 10, covering model development, model evaluations, attacks and failure modes, agency, deployment, cybersecurity and impacts. Each item is labelled "Risk management measure" (or "Risk management measure (Experimental)") or "Risk source", followed by descriptive text. The work is deliberately descriptive rather than prescriptive so that text can be inserted directly into technical standards and codes of practice, and the authors state that Section 3 and the catalog are placed in the public domain under CC0 1.0.

The conclusion discusses false negatives in risk assessment, the difficulty of allocating resources for risk management and "safetywashing", and considers small and medium enterprise GPAI providers separately.

🧩 What’s Covered

  • Terminology and item formatting (Section 2): defines GPAI model, GPAI system, GPAI provider, systemic risk, risk, harm, risk source and risk management measure; notes that GPAI models trained with more than 10^25 FLOPs are presumed to pose systemic risk under the EU AI Act, and fixes the item format of type, title and description.
  • Safety engineering process (Section 3): places the catalog inside an iterative safety engineering loop (Figure 1), stating that the risk source content supports identification step (B), that the measure content supports steps (C), (G) and (H), and that the list is designed to be used as a checklist.
  • Model development (Section 4): data-, training- and fine-tuning-related items, including difficulty filtering large web scrapes, missing cross-organizational documentation, adversarial examples and robust overfitting, alongside measures such as synthetic data, adversarial training, data cleaning and tamper-resistant safeguards for open-weight models.
  • Model evaluations (Section 5): general evaluations, benchmarking with benchmark leakage, raw data, cross-lingual, guideline, annotation and post-deployment contamination, red teaming, auditing and interpretability/explainability, including auditor conflicts of interest and auditor capacity mismatch.
  • Attacks and failure modes (Section 6): jailbreaks, multimodal jailbreaks, transferable adversarial attacks, backdoors, text-encoding attacks, many-shot jailbreaking, distraction by irrelevant context and knowledge conflicts in retrieval-augmented LLMs.
  • Agency (Section 7): goal-directedness (specification gaming, measurement and reward tampering, goal misgeneralization), deception, situational awareness, self-proliferation and persuasion.
  • Deployment (Section 8): risk assessment techniques drawn from other safety-critical industries (margin of safety, scenario analysis, fishbone diagram, causal mapping, Delphi technique, cross-impact analysis, bow-tie analysis, STPA, risk matrices), staged and restricted model release, post-deployment practices such as pop-up interventions, watermarking and metadata, and monitoring.
  • Cybersecurity, impacts, discussion and appendices (Sections 9–12 and Appendices A–D): least privilege, sandboxing, structured access and bandwidth limits; physical, societal, financial, cyberattack, weapons, bias, privacy and environmental impacts; SME considerations; and background on the GPAI value chain, benchmarking risks, autonomous AI systems and a proposed risk taxonomy.

💡 Why it matters?

Providers, standards writers and auditors can use the catalog as a structured checklist when scoping risk assessments, because each entry separates "what is it?" from "when to use it?" and can be copied directly into codes of practice under the public domain licence. The evaluation and auditing items help teams design red-teaming, benchmarking and third-party access arrangements, while the deployment section imports techniques proven in other safety-critical industries. The paper states its alignment with EU AI Act terminology and lists ISO/IEC 42001, ISO/IEC 23894, ISO/IEC 27001 and ISO/IEC 27002 as standards whose measures it largely avoids repeating.

❓ What’s Missing

The authors state that the catalog is not exhaustive and that any derivative should not be assumed exhaustive, and it offers no prioritisation: risks are not weighted by likelihood or severity, which is left to "the relevant actors". Measures are descriptive only, with no criteria for when each should be applied, and many are marked experimental with little mature knowledge of their drawbacks. Cost, staffing and implementation detail are largely absent beyond pre-allocating resources, and no legal obligations are mapped. Regulatory examples are mostly EU AI Act-based, and the available extraction covered 80 of 92 pages, ending partway through Appendix A.

👥 Best For

Standards experts and policymakers drafting GPAI safety standards and codes of practice; provider model safety and system safety engineering teams building risk checklists for development, evaluation and release; risk management, audit and red-teaming functions designing evaluation and third-party access arrangements; and researchers mapping GPAI risk sources to mitigations.

📄 Source Details

Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems is an arXiv preprint (arXiv:2410.23472v2, 15 November 2024) by Rokas Gipiškis, Ayrton San Joaquin, Ze Shen Chin, Adrian Regenfuß, Ariel Gil and Koen Holtman, affiliated with the AI Standards Lab, Vilnius University and Holtman Systems Research. Section 3 and the catalog are placed in the public domain under CC0 1.0 and the rest submitted under CC BY 4.0. The document runs 92 pages; the available extraction covered 80 pages, ending partway through Appendix A.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.