AI Governance Library

AI Safety Index

An independent assessment of seven leading AI companies’ practices for managing immediate harms and catastrophic risks from advanced AI. It grades 33 indicators across six domains and sets out company-specific improvement opportunities.
Cover of AI Safety Index

⚡ Quick Summary

Published by the Future of Life Institute, this report presents the Summer 2025 edition of the AI Safety Index: an independent assessment of seven frontier AI developers’ management of immediate harms and catastrophic risks from advanced AI. It assesses Anthropic, OpenAI, Google DeepMind, Meta, xAI, Zhipu AI and DeepSeek, using evidence gathered from 24 March to 24 June 2025 and reviewed by an independent expert panel. The Index is intended as a public-facing tool for comparing corporate safeguards, tracking behaviour and identifying gaps between safety commitments and implemented practice.

The Index grades companies on 33 indicators across six domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing. It uses domain-level A–F grades based on absolute performance standards, with final scores averaging reviewer grades. Anthropic receives the highest overall result, C+ (2.64), followed by OpenAI at C (2.10) and Google DeepMind at C- (1.76); no company scores above D in Existential Safety. The report’s central finding is that companies’ capability ambitions are advancing faster than their risk-management practice, and that none has presented what reviewers considered a coherent, actionable plan for controlling AGI or more capable systems.

🧩 What’s Covered

The report proceeds from its rationale and methodology to comparative grades, domain findings and supporting grading materials.

  • Purpose, scope and rankings: Introduces the Index as an assessment of safeguards around increasingly capable general-purpose AI systems. It presents overall grades and numerical scores for all seven companies, alongside the report’s key findings and suggested improvement opportunities for each firm.
  • Company selection and index design: Explains that developers were selected for deployment of models with competitive public-benchmark performance. The 33 indicators are selected for signal value, implementation focus, information availability, clear definition and recognition of leadership.
  • Evidence collection and review: Describes desk research using public technical, policy and company materials, external benchmarks, media reporting and a 34-question company survey. Three companies—OpenAI, Zhipu AI and xAI—responded. Six independent experts assigned domain grades, with no fixed weighting imposed across indicators inside a domain.
  • Risk Assessment and Current Harms: Covers dangerous-capability evaluations, elicitation, human uplift trials, independent review, pre-deployment testing and bug bounties. It also compares benchmark and adversarial-testing results, fine-tuning safeguards, watermarking and default treatment of user inputs.
  • Safety Frameworks and Existential Safety: Assesses published frontier safety frameworks through risk identification, analysis, treatment and governance, incorporating SaferAI’s analysis. It also examines strategies for alignment and control, internal monitoring, technical safety research and support for external safety researchers.
  • Governance, accountability and information sharing: Reviews lobbying positions, company structures, whistleblowing-policy transparency and reporting culture. It assesses system-prompt and behavioural-specification disclosure, participation in the G7 Hiroshima AI Process and the Index survey, incident reporting and public communication about extreme risks.
  • Limitations and appendices: Sets out constraints arising from reliance on public disclosures, inability to independently verify company claims, uneven information availability and Western-centric assumptions. Appendix A supplies domain grading sheets, while Appendix B contains the company survey.

💡 Why it matters?

For AI governance, risk and assurance teams, the Index makes otherwise dispersed evidence about frontier developers more comparable. Its indicators connect organisational choices—such as whistleblowing arrangements, external testing access, incident reporting and risk-governance structures—to technical evaluation and deployment practices. The report also distinguishes stated commitments from implemented assessments, an important distinction for those interpreting company safety claims.

The findings identify practical weaknesses that affect external scrutiny: the panel found no company had commissioned an independent verification of its internal safety evaluations, and none had a concrete public process for notifying governments about critical incidents. It also relates framework assessment to measurable thresholds, mitigation measures and oversight arrangements.

❓ What’s Missing

The report explicitly cautions that it is not a comprehensive assessment of all AI-safety dimensions. Its reliance primarily on public information means low transparency can be difficult to distinguish from poor implementation, and official company claims cannot be independently verified. It cannot assess some critical practices, including cybersecurity investment to protect model weights, because public disclosure is limited. The 33 indicators cover only practices for which meaningful evidence was available. The authors also acknowledge that the methodology was developed in Western academic contexts and may disadvantage Chinese companies, particularly where it values self-governance and information sharing that operate differently across regulatory and cultural settings. Finally, flexible weighting by reviewers and a six-person panel may introduce inconsistency or leave relevant expertise unrepresented.

👥 Best For

Useful for AI governance and risk leads comparing frontier-developer disclosures; policy teams examining corporate positions on safety regulation; evaluators designing dangerous-capability or external-testing processes; and boards or researchers assessing how governance, whistleblowing and transparency practices interact with technical safety claims. The grading sheets are particularly useful for readers seeking indicator-level evidence and assessment criteria.

📄 Source Details

AI Safety Index is an English-language Summer 2025 publication of the Future of Life Institute, dated 17th July 2025. No individual authors are named. The document states that it is available online at futureoflife.org/index. The supplied input is a partial extraction: 80 of 101 PDF pages were available, covering PDF pages 1–80; this review describes only that available material.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.