AI Governance Library

A Formal Model of How AI Erodes Human Agency

In this report, we develop a formal framework for measuring how artificial intelligence (AI) affects collective human agency in decisionmaking processes. We define collective agency as humanity’s capacity to make consequential choices about shared resources, institutions, and the social environment.
A Formal Model of How AI Erodes Human Agency

⚡ Quick Summary

Published by the RAND Corporation, this report introduces a formal, model-theoretic framework grounded in social choice theory to quantify how artificial intelligence alters collective human agency in decision-making processes. Defining collective agency as humanity's capacity to make consequential choices regarding shared resources, institutions, and social environments, the authors model decision-making through decisive coalitions—subgroups capable of determining aggregate outcomes. The report identifies and mathematically models three distinct erosion mechanisms: human disenfranchisement, AI enfranchisement, and AI agenda control. Furthermore, it defines a formal terminal state of agency loss characterized by a single minimal decisive coalition, providing mathematical metrics to track power concentration and evaluate early tipping points where agency loss becomes irreversible.

🧩 What's Covered

The report establishes a coalition-based framework to analyze structural shifts in collective decision-making across several core areas:

  • Coalition-Based Agency Metrics: Defines three quantitative indicators: the per-capita distribution of decisive coalitions w(G), the relative size of the smallest decisive coalition m(G), and the composition of minimal decisive coalitions f(G), leveraging filter and ultrafilter algebraic structures from social choice literature.
  • Irreversibility Dynamics: Applies economic definitions of irreversibility to AI adoption, noting that labor substitution, cognitive atrophy, and dismantled deliberative infrastructure make recovering lost human decision-making capacity prohibitively costly.
  • Three Mechanisms of Agency Loss: Formally validates three distinct pathways:
    • Human Disenfranchisement: Removing humans from decision-making roles, causing average per-capita agency to vanish as disenfranchised populations increase.
    • AI Enfranchisement: Direct substitution of humans by AI entities, reducing the probability that outcomes are determined exclusively by human consensus even when coalition sizes remain constant.
    • AI Agenda Control: Autonomous AI entities introducing or manipulating alternative choice menus to split human preferences and force aggregate outcomes aligning with AI preferences.
  • The Terminal End State: Formally characterizes the theoretical extreme of agency decay, wherein decision-making power consolidates into a single minimally sized decisive group.
  • Practical Applications and Recommendations: Outlines use cases for evaluating human-in-the-loop systems (such as command-and-control OODA loops in autonomous weapons) and developing AI agency benchmarks. It calls on policymakers to establish minimum human participation thresholds in decisive coalitions for high-stakes domains and benchmark reversibility capacity.

💡 Why it matters?

Most AI safety and governance frameworks focus exclusively on model capabilities, alignment, or immediate security risks, overlooking gradual, structural disempowerment. This report demonstrates that even perfectly aligned AI systems that faithfully produce outcomes humans endorse can permanently erode human agency by displacing active participation. By translating social choice theory into quantifiable metrics, the report equips governance, risk, and compliance leaders with analytical tools to detect nonlinear accelerations in power concentration before institutional knowledge and oversight capacity degrade beyond recovery.

❓ What's Missing

The authors explicitly acknowledge several modeling boundaries:

  • Narrow Task Scope: Focuses primarily on choice aggregation tasks and menu modification rather than broader decision-structuring tasks, such as initial problem formulation and objective setting.
  • Binary Decisiveness: Models coalition participation as binary influence, omitting qualitative degradation like cognitive atrophy where humans retain formal voting power but lack substantive judgment.
  • Single-User Dynamics: Excludes individual chatbot-human advisory interactions that do not involve preference aggregation across groups.

👥 Best For

This report is best suited for AI governance researchers, safety evaluators, public policymakers, national security strategists, and enterprise risk officers designing oversight frameworks, benchmarks, or human-in-the-loop requirements for autonomous decision systems.

📄 Source Details

  • Title: A Formal Model of How AI Erodes Human Agency
  • Authors: Alvin Moon and Benjamin Boudreaux
  • Organization: RAND Corporation (Center for the Geopolitics of Artificial General Intelligence)
  • Publication Date: 2026
  • Report ID: RR-A4817-1
  • Resource Type: Research Report
  • Source Link: rand.org/t/RRA4817-1

📝 Thanks to

Thanks to Alvin Moon and Benjamin Boudreaux for authoring this research, as well as the RAND Center for the Geopolitics of Artificial General Intelligence and reviewers Robert Lempert and Raymond Douglas for their contributions to formalizing AI agency measurement.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.