Anthropic’s Use of Copyrighted Books for AI Training Deemed Fair Use by US Court
In a closely watched ruling, the US District Court for the Northern District of California found that training large language models on lawfully acquired copyrighted books is transformative fair use. The court also held that converting purchased print books to digital format for training purposes is protected, but it sharply distinguished the use of pirated materials as inherently infringing.
The case, Andrea Bartz et al. v. Anthropic PBC, Case No. 24-CV-05417-WHA (N.D. Cal. June 23, 2025) (Alsup, J.), examined Anthropic’s practices around compiling a massive central digital library and using it to train its AI assistant, Claude. Authors Andrea Bartz and two others sued after their works were copied – both from purchased print editions that were digitized and from over seven million books downloaded from pirate sites without permission.
The legal bedrock for the decision is Section 107 of the Copyright Act, the fair use doctrine:
Notwithstanding the provisions of sections 106 and 106A, the fair use of a copyrighted work, including such use by reproduction in copies or phonorecords or by any other means specified by that section, for purposes such as criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research, is not an infringement of copyright. In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include—
(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
(2) the nature of the copyrighted work;
(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
(4) the effect of the use upon the potential market for or value of the copyrighted work.
The fact that a work is unpublished shall not itself bar a finding of fair use if such finding is made upon consideration of all the above factors.
The district court broke Anthropic’s actions into three distinct categories.
Transformative training (fair use)
The authors challenged the inputs used to train the LLMs, not their outputs. The court found that using copyrighted books as training data was transformative, comparing it to how humans read and learn from texts to write new, original material. No evidence showed that Claude publicly released infringing copies. The training process – enabling the generation of new content rather than direct replication – weighed heavily in favor of fair use.
Format-shifting copies (fair use)
Anthropic lawfully purchased print editions, removed bindings, scanned pages, and stored digital copies, destroying the originals. The authors did not allege that the company distributed these digital copies outside Anthropic. The court held that the print-to-digital conversion, done to save space and enable efficient search, was a transformative fair use.
Liability for piracy (not fair use)
The court drew a firm line against the downloading and retention of more than seven million books from pirate websites, regardless of later training use. Even after Anthropic opted not to train on those copies, it kept them as part of a central research library – an act the court found inherently infringing and non-transformative. The ruling stressed that each act of copying must be judged by its own objective use, and that “piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.” Anthropic now faces a jury trial limited to damages for those pirated copies.
Practice note: This is the first federal court decision analyzing fair use of copyrighted material to train generative AI. Two days later, another judge in the same district ruled in _Kadrey et al. v. Meta Platforms Inc._ that Meta’s use was also transformative, though he based his ruling on the plaintiffs’ failure to show market impact.
—
The ruling is a significant step toward clarifying how copyright law applies to AI training. Andrew Ng, writing in The Batch, called it “reasonable and good for AI progress,” noting that training on legitimately acquired data to produce transformational outputs is now explicitly endorsed.
He described how data preparation dominates the day-to-day work of foundation model teams – cleaning, tokenizing, and error-analyzing high-quality sources like books – and observed that the decision removes a major risk to data access. At the same time, the court’s refusal to extend fair use to pirated materials will force providers to scrutinize datasets that may contain unlicensed works.
Ng acknowledged the anxiety felt by writers about their livelihoods but emphasized that clear legal guardrails can help the industry move forward while leaving room for fair compensation mechanisms down the road.
Sources:
