Press Release

Book, News, and Journal Publishers File Amicus Brief in In Re Mosaic LLM Litigation

Book, News, and Journal Publishers File Amicus Brief in In Re Mosaic LLM Litigation

Today, the Association of American Publishers (AAP), News/Media Alliance (N/MA), and International Association of Scientific, Technical & Medical Publishers (STM) filed an amicus brief in In Re Mosaic LLM Litigation. This case was brought in March 2024 by a group of authors alleging, among other things, that the defendants, Mosaic and Databricks, downloaded datasets including pirated copies of copyrighted textual works, and trained large language models (LLMs) using copyrighted works.  The case is before Judge Charles Breyer in the Northern District of California.

The brief urges the district court to deny the defendants’ motion for summary judgment and find that their unauthorized use of copyrighted works for AI training is not fair use.  As the brief explains, fair use is a critical exception to exclusive rights that should not be extended to LLMs that exploit expressive content to generate expressive content, essentially serving substitutes for copyrighted works.  Unlicensed AI training substantially and adversely usurp markets that rightfully belong to copyright owners, including markets for derivative works and AI licensing. 

Excerpts from the brief:

  • Overly expansive interpretations of fair use displace the market mechanisms Congress established, replacing negotiated exchange with uncompensated (and uncredited) use and transferring value from the authors and publishers that invest in them to industries that neither incurred the costs nor assumed the risks of producing the underlying works.
  • When properly analyzed, LLMs, like those of the defendants, exploit copyrighted works for their expressive value for the same purpose, ultimately generating substitutes for the works used for training.  Fair use should not sanction AI systems with an inherent “problem of substitution.”
  • Consistent with their purpose, design, and training, LLMs can readily generate substitutes for the copyrighted works on which they are trained, resulting in non-transformative uses.  Those substitutes can take the form of verbatim and near-verbatim copies, summaries and alternative versions of written works, knock-offs that copy expression and creative choices from original works, and derivative works exclusively reserved for rightsholders.  LLMs’ proven ability to generate outputs that substitute for, are derivatives of, or otherwise exploit the expressive content of the original works further demonstrates that LLMs are not highly transformative. 
  • Consumers are confused and overwhelmed, and copyright owners are losing revenues and market visibility to the flood of AI-generated books and other text.
  • [L]icensing provides AI developers authorized access to human-created works and high-quality content.  Using unauthorized content from online or questionable sources carries the risk of tainting training sets with low-quality, such as AI-generated, text.  For AI to flourish, licensing is needed to ensure humans continue to be incentivized to create so that the well of human creation does not run dry and that AI companies have reliable access to content that improves model performance.

The full amicus brief is available here.