contact us
A status check on the major copyright lawsuits against AI companies and the legal questions they raise.
Copyright Lawsuits Against AI Companies: Where the Key Cases Stand

Since the current generation of generative AI tools reached wide adoption, rights holders (authors, news publishers, visual artists, musicians, and their respective industry organizations) have filed numerous lawsuits against AI companies over the use of copyrighted material in training data and, in some cases, over the copyright status of AI-generated output itself. These cases are working through courts in multiple jurisdictions, and while outcomes remain genuinely unresolved as of this writing, the core legal questions being litigated are worth understanding clearly, independent of any single case's outcome.

The central question: does training on copyrighted material infringe?

The foundational legal question across most of these cases is whether using copyrighted text, images, or audio to train an AI model constitutes copyright infringement, or whether it falls under fair use (in US law) or an equivalent exception in other jurisdictions. AI companies have generally argued that training constitutes a transformative use, the model isn't reproducing or distributing the copyrighted work itself, but learning statistical patterns from it, analogous to how a human learns writing style from reading widely without each act of reading constituting infringement. Rights holders have generally argued that large-scale, systematic copying of copyrighted works for commercial model training is different in kind from individual human learning, and that it causes direct economic harm by enabling a competing product built substantially on their work, often without compensation or consent.

A second, distinct question: does output infringe?

Separate from the training-data question, some cases center on whether specific AI-generated outputs infringe copyright. For instance, when a model can be prompted to reproduce substantial, recognizable portions of a specific copyrighted work, or when generated content is substantially similar to a specific existing work in ways that go beyond general stylistic influence. This is a more traditional copyright question in some ways, but it intersects with the training-data question because plaintiffs often argue that a model's capacity to reproduce protected material is itself evidence that infringing copying occurred during training, not just at generation time.

Why licensing deals are happening alongside litigation

Running in parallel with contested litigation, several AI companies have negotiated licensing agreements with publishers and content owners. Paying for the right to train on specific, clearly licensed content and, in some deals, for ongoing access to current content. These deals don't resolve the broader legal questions being litigated elsewhere, but they represent a pragmatic middle path some rights holders and AI companies have found preferable to years of uncertain litigation, and they're likely to become more common regardless of how the contested cases are ultimately decided.

Why the outcome matters beyond the named parties

However these cases resolve, the precedent will shape the AI industry's data practices well beyond the specific companies and rights holders involved, a ruling that broadly favors AI companies would likely reinforce current training practices industry-wide; a ruling favoring rights holders could force significant changes to how models are trained, potentially including retroactive liability questions for models already trained and deployed. This uncertainty is itself a live business risk that shows up directly in AI company valuations and in enterprise buyers' due diligence, connecting to the deal-diligence questions we cover in our coverage of the AI acquihire boom, where training-data provenance has become a standard diligence item.

How this intersects with regulation

Separately from litigation, some regulatory frameworks have begun addressing training-data transparency directly, the EU AI Act, covered in our explainer, includes disclosure obligations for general-purpose model providers regarding training data at a summary level, which is a regulatory approach to transparency running alongside, rather than replacing, the ongoing copyright litigation in courts.

The honest state of play

These are live, contested, unresolved legal questions, not settled ones, and any specific case's outcome (including on appeal) could meaningfully shift the broader legal landscape. Treat any confident claim about "how this will turn out" with appropriate skepticism; the honest current state is genuine legal uncertainty.

Share with