AI Governance Library

Copyright and Artificial Intelligence, Part 3: Generative AI Training

U.S. Copyright Office report, Part 3 of its Copyright and Artificial Intelligence study, analysing how copyrighted works are used in generative AI training and how fair use and licensing apply, released in a pre-publication version.
Cover of Copyright and Artificial Intelligence, Part 3: Generative AI Training

⚡ Quick Summary

Published by the U.S. Copyright Office, this is Part 3 of its Copyright and Artificial Intelligence report, released in a pre-publication version in May 2025. It addresses the use of copyrighted works in developing generative AI systems: whether those acts require owners' consent or compensation, and how that could feasibly be accomplished.

The available text covers a technical overview of machine learning, training data, training phases, memorization and deployment; a prima facie infringement analysis of data collection, training, retrieval-augmented generation and outputs; and a factor-by-factor fair use analysis under Section 107, including transformativeness, commerciality, unlawful access, market dilution and lost licensing opportunities. The Office concludes that training a foundation model on a large and diverse dataset will often be transformative, that different uses during development and deployment require separate consideration, and that some uses will qualify as fair use while others will not. Its table of contents lists a further section on licensing for AI training and a conclusion, which fell outside the extracted pages.

🧩 What’s Covered

  • Technical background (Section II): how generative models are built, covering statistical models and neural network weights, generative pre-training through next-token prediction, tokenization and context windows, data characteristics of quantity, quality and purpose, acquisition through scraping, licensing and pirate sources, and curation by filtering, cleaning and compiling.
  • Training and deployment: training phases labelled "pre-training," "post-training" and "fine-tuning" and why those labels are misleading; the disputed extent of memorization and the factors influencing it; and deployment through retrieval-augmented generation, guardrails, alignment training and varying developer control over model weights.
  • Prima facie infringement (Section III): reproduction and derivative work rights across data collection and curation, training, RAG and outputs, including discussion of Kadrey v. Meta Platforms and Andersen v. Stability AI on whether model weights infringe.
  • Fair use factor one (Section IV.A): identifying the specific use, transformativeness after Warhol, the "non-expressive use" and human-learning arguments the Office rejects, commerciality and "data laundering," and unlawful access to pirated or paywalled works.
  • Factors two, three and four: the nature of the copyrighted work; the amount used, its reasonableness in light of purpose and the amount made available to the public; and market harm through lost sales, market dilution, lost licensing opportunities and public benefits.
  • Weighing, competition and international approaches: the first and fourth factors assuming considerable weight, a spectrum from likely-fair research uses to unlikely-fair training on pirate sources, competition concerns referred to antitrust law, and text and data mining exceptions in the EU, Japan and Singapore.

💡 Why it matters?

For anyone assessing how copyright applies to AI development, the report supplies a structured analytical framework rather than a verdict: it separates the distinct acts of dataset compilation, pre-training, fine-tuning, RAG and output generation, and applies each Section 107 factor to those acts. That separation is directly useful to legal and compliance teams mapping exposure, because the Office states that different uses require separate consideration and that commerciality turns on whether a specific use furthers commercial purposes rather than on an entity's status. It also documents the fact patterns courts are already testing, including piracy, paywall circumvention and the efficacy of guardrails.

❓ What’s Missing

The report states its analysis is "necessarily limited to current circumstances and publicly available information," and the Office says it will monitor developments and may revisit conclusions. It offers no opinion on specific pending cases and no legislative text. The extraction available here ends within the international approaches discussion, so the licensing analysis and conclusion, listed only in the table of contents, could not be reviewed. Several questions are left open on the document's own account: how much memorization occurs, whether guardrails are effective, and whether licensing markets can exist for works whose ownership is diffuse.

👥 Best For

Legal and policy teams tracking how U.S. copyright law applies to AI training; compliance leads mapping where copying occurs across the AI development pipeline; and standards, audit or risk specialists who need the Office's factor-by-factor reasoning on transformativeness, memorization, guardrails and market dilution as a reference point.

📄 Source Details

The resource is Copyright and Artificial Intelligence, Part 3: Generative AI Training, a pre-publication version of a report of the Register of Copyrights, published by the U.S. Copyright Office in May 2025. No individual authors are named; it is issued in the Office's name. The PDF runs to 113 pages and carries no series number beyond the Part 3 designation, and no URL for the document itself is printed in the extracted text. The input was a partial text extraction covering 80 of the 113 pages, from the cover through Section IV.G.

About the author
Jakub Szarmach

AI Governance Library

Curated Library of AI Governance Resources

AI Governance Library

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to AI Governance Library.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.