The Underappreciated Third U.S. Fair Use Factor in Copyright Infringement Cases Concerning AI Training
August 5, 2026
In many copyright infringement cases filed in the United States against AI companies because of their unlicensed use of copyrighted works in the training of AI models, defendants raise fair use as a defense to the infringement claims. It is presently unclear whether U.S. appellate courts will find the use of copyrighted works to train AI to be fair use. Although supporters of the AI companies claim that the outcomes of the existing three district court decisions are clearly in favor of fair use in the AI training cases, the outcomes are by no means unequivocal.
In the continuing cross-fire of arguments, defendants and plaintiffs in these cases have much to say about fair use but, of the four factors in the non-exhaustive list of fair use factors in the U.S. Copyright Act, the third factor, “the amount and substantiality of the portion used in relation to the copyrighted work as a whole,” appears to be woefully underappreciated in the cases. This factor deserves much more attention than it is receiving, and in these cases, it should be the key factor in the resolution of the fair use analysis.
The first and the fourth fair use factors draw the most attention in generative AI training cases. “The purpose and character of the use” and “the effect of the use upon the potential market for or value of the copyrighted work,” respectively, have played key roles and driven discussions about whether the use of copyrighted works to train AI can be described as transformative, and whether AI output can be understood as a substitution—even if as an indirect substitution—for the copyrighted works that are used in the training process.
The analysis of the third factor has been subordinated to the consideration of the first factor, the purpose and character of the use; the first factor drives how much and what can be used from a work to remain within fair use. The end thus justifies the means, but the U.S. Supreme Court made clear in Campbell that fair use requires that no more of a work be taken than necessary to achieve the purpose of the use.
In the generative AI training cases, defendants have argued that entire works are required to train AI; the Register of Copyrights’ Report on Generative AI Training, which was pre-published in May 2025, recognized that “the use of entire works appears to be practically necessary for some forms of training for many generative AI models.” The use of entire works would not necessarily weigh against fair use; indeed, in some previous non-AI cases, courts found that even when entire copyrighted works were used, the use could still be fair.
In these previous cases, the use of entire works did not weigh against fair use because such use was reasonably necessary for the purpose to be achieved. In Sony v. Universal, a majority of the U.S. Supreme Court decided that the reproduction of entire copyrighted works for purposes of time-shifting did not weigh against fair use; in Authors Guild v. HathiTrust, the U.S. Court of Appeals for the Second Circuit found that the use of entire works was not excessive for the full-text search function; in Kelly v. Arriba and Perfect 10 v. Amazon.com, the U.S. Court of Appeals for the Ninth Circuit found that the third factor weighed neither for nor against the defendants when the defendants copied entire copyrighted works in order to display them as thumbnail images in search engine results; and in A.V. v. iParadigms, the U.S. Court of Appeals for the Fourth Circuit agreed with the lower court that the use of entire copyrighted works for a plagiarism check system was sufficiently “limited in purpose and scope” and justified by the transformative nature of the use.
The important point in all these cases, however, is that the use of the entire copyrighted works did not weigh against fair use because the purpose of the use also required the use of the particular works; the purpose in these cases could not have been achieved without the specific works being used. The majority in Sony v. Universal conditioned its decision regarding the third factor on the fact that the “timeshifting [in the case] merely enable[d] a viewer to see such a work which he had been invited to witness in its entirety free of charge” (emphasis added). In Authors Guild v. HathiTrust, the court found that “the record demonstrate[d] that these copies are reasonably necessary to facilitate the service” (emphasis added). In Kelly v. Arriba the court noted that “[i]t was necessary for Arriba to copy the [particular] entire image to allow users to recognize the image and decide whether to pursue more information about the image or the originating web site.”
The problem in generative AI training cases is that no one particular work seems to be required to train an AI model; while in the abstract, AI training might require a vast number of entire works, in the concrete, the training does not justify the use of any particular work—unless the AI is trained to be able to regurgitate the specific copyrighted works used for the training, which ability the AI companies deny. Judge Bibas pointed out this issue in his February 2025 revised opinion in Ross, a case that involved non-generative AI. In the analysis of the first fair use factor, the judge noted that while the copying in the earlier cases concerning reverse engineering (Sega v. Accolade and Sony Comp. v. Connectix) “was necessary for competitors to innovate,” in the case before him, the “copying [was] not reasonably necessary to achieve the user’s new purpose” (quoting Warhol). Thus in Concord Music v. Anthropic, Concord Music et al. suggested in their filing against Anthropic that AI companies would have to show a “specific need for [the copyrighted] works or that alternatives were unavailable” (Plaintiffs’ Notice of Motion and Motion for Preliminary Injunction, document 179, August 1, 2024, p. 29).
If a particular work is so crucial to AI development that the AI model could not exist without that work in its training data, then the resulting indispensability might justify the use of the entire work for the training. However, such an indispensability of the work might in turn raise suspicions about what content from the work is truly indispensable—the non-copyrightable elements, or a set of elements that are copyrightable in their sum (see also Concord Music et al. here).
The fine-tuning stage of AI training might demonstrate when and why particular works are required for some stages of the training. As three academics–amici curiae pointed out in Ross, “because fine-tuning requires specialized knowledge, developers typically rely on subject-matter experts to handcraft training examples;” the “experts directly enhance pretrained models by conditioning them to produce reliable responses about specialized domains.” But while fine-tuning might be a special case, the process of pre-training, or general AI training, does not justify the use of any particular copyrighted work.
The Register of Copyrights’ Report acknowledged that “for large, general-purpose models, there is no need to copy any amount of any specific work.” If, as the Supreme Court said in Campbell, “the amount and substantiality of the portion used in relation to the copyrighted work as a whole [must be] reasonable in relation to the purpose of the copying,” then, if the purpose does not necessitate the use of a particular work at all, the use of any portion of a work must be unreasonably large for the purpose. No large degree of transformativeness of the use of a work and no minimal degree of substitution of the output for an original work should outweigh the fact that the particular work is not necessary for the use. Consequently, plaintiffs should prevail on the strength of the third factor, which in these cases does not only “strongly weigh” against fair use, as plaintiffs have argued, but completely outweighs all other factors. Cumulatively, AI training requires the use of many “diverse sources,” but this requirement does not justify the use of any one particular copyrighted work.
The fact that AI training cases are often brought as class action lawsuits can complicate the analysis of the third factor. Class action litigation presents some significant advantages, but for the third fair use factor a class action might create problems because it is more difficult to argue that an entire class of works is not necessary for AI training. Nevertheless, class action conditions of commonality and typicality would seem to be satisfied, even when responses to a fair use defense are based on a lack of necessity of use of individual copyrighted works.
The author thanks for their assistance in researching AI training cases her colleagues at the Wiener-Rogers Law Library at the William S. Boyd School of Law, including Research Librarian and Associate Professor Youngwoo Ban, Assistant Professor Tina Fortier, and Library law student research assistants David (Job) Dooley (‘26) and Miranda Romero (‘27). For another output of the research, see a forthcoming chapter.