AI Governance Library

Classifying AI Systems

This CSET Data Brief evaluates four frameworks for classifying AI systems along policy-relevant characteristics. It reports survey experiments testing the consistency and accuracy of classifications by public users.
Cover of Classifying AI Systems

⚡ Quick Summary

Published by the Center for Security and Emerging Technology, this report examines how classification frameworks can support AI governance by characterising systems through observable, policy-relevant attributes. It argues that a shared classification approach can help developers, governing bodies, and users compare systems, monitor risk and bias, manage inventories, and make more targeted regulatory decisions. The work was developed through discussions with the OECD AI Policy Observatory and the U.S. Department of Homeland Security Office of Strategy, Policy, and Plans.

The brief evaluates four frameworks based on the OECD definition of an AI system. Frameworks A, B, and C classify systems by autonomy and impact, while Framework D covers context, input, model, and output through nine dimensions. CSET tested the frameworks in two survey rounds in which 361 respondents completed 1,831 classifications. The findings show that Framework C's descriptive autonomy labels and the use of summary rubrics improved results; respondents classified impact more reliably than autonomy. The brief concludes that usable classification also depends on access to adequate information about a system, especially information on technical characteristics.

🧩 What’s Covered

The brief proceeds from the case for classification to framework development, testing, and future work.

  • Governance rationale and AlphaGo Zero example: Explains classification as assigning predefined labels to observable characteristics, rather than pursuing a single fixed definition of AI. AlphaGo Zero is used to show how an AI system can be classified as high autonomy and low impact, compared with other systems, and associated with a risk profile and proportionate governance needs.
  • Existing classification approaches: Reviews technical frameworks based on algorithmic openness, data, AI approaches, and problem domains, alongside approaches centred on automation. It contrasts these with policy interest in human interaction, impact, risk, and rights, well-being, and freedoms.
  • Framework development: Describes discussions with DHS and the OECD Network of Experts on AI working group. Frameworks A–C use autonomy and impact, with different autonomy labels; Framework D covers sector, impact, criticality, system user, data collection, data structure, acquisition of capabilities, task, and autonomy.
  • Survey methodology: Sets out two survey rounds using short descriptions of deployed systems including missile defence, drug-interaction prediction, credit scoring, facial-image quality, and search and rescue. Respondents recruited through Mechanical Turk were randomly assigned frameworks, and consistency and accuracy were assessed using a 65 percent threshold.
  • Results for Frameworks A, B, and C: Reports that Framework C produced 44 percent accurate and consistent system-dimension classifications. Its action, decision, and perception labels were more descriptive than the high, medium, low, and none labels in Framework A; a reference rubric also improved performance.
  • Results for Framework D and cross-framework comparisons: Finds 51 percent consistency across Framework D's dimensions, with stronger results for context and task than for input, model, or autonomy. Comparisons restricted to autonomy and impact found near-identical consistency for Frameworks C and D, while Framework C performed best for the five systems classified under every framework.
  • Discussion and next steps: Identifies accessible system information as central to classification, notes difficulty with technical characteristics, and outlines further OECD consultation analysis, incident-based system identification, and possible automated extraction from public text.

💡 Why it matters?

For AI governance teams, the brief treats classification as a practical precursor to risk assessment, inventory management, and targeted oversight. It shows that dimensions such as impact, deployment context, task, and autonomy can structure information about varied systems without treating all AI systems as equivalent.

The experiment also identifies implementation constraints. People were more successful at classifying impact than autonomy, and technical dimensions were harder to assess when descriptions lacked sufficient detail. Teams using such a framework therefore need clear rubrics and accessible evidence about a system's use case, inputs, model capabilities, and human involvement rather than relying only on high-level system descriptions.

❓ What’s Missing

The brief tests four candidate frameworks rather than establishing a single final framework for adoption. Its principal survey sample was recruited through Amazon Mechanical Turk and restricted to respondents located in the United States with specified education and platform-qualification criteria; only 11 U.S. government personnel responded, so their findings were not reported in the main analysis. The assessment relies on short descriptions based on publicly available information and a limited set of example systems. The author also notes that expert classifications involved disagreement and incomplete knowledge of the systems, making accuracy a weaker performance measure. The expanded OECD framework and public-consultation responses were still under analysis at publication.

👥 Best For

Policy teams building AI inventories, risk-assessment processes, or system-characterisation practices will find the framework dimensions and survey findings useful. It is also suited to researchers and programme leads comparing alternative ways to describe autonomy, impact, deployment context, inputs, models, and outputs, especially where classifications must be understandable to non-specialist users.

📄 Source Details

Classifying AI Systems is an English CSET Data Brief by Catherine Aiken, published by the Center for Security and Emerging Technology in November 2021. The supplied input is the complete 51-page PDF. It carries the document identifier doi: 10.51593/20200025 and a Creative Commons Attribution-Non Commercial 4.0 International License.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.