Sony Music Publishing and Warner Chappell filed their new case against Anthropic on August 28 in federal court in Northern California. Anthropic co-founders Dario Amodei and Benjamin Mann are also named as individual defendants.
The publishers allege that tens of thousands of copyrighted compositions were copied without licenses as part of Claude's development.
The acquisition layer may be the most important part of the case
Much of the AI copyright debate starts with training itself. A model processes a protected work, changes internal parameters and eventually produces new outputs. AI companies argue that this can qualify as a transformative use.
The new complaint attacks an earlier step.
Sony and Warner allege that Anthropic obtained some copyrighted material through scraping, downloads, torrents and sources the publishers characterize as pirate repositories.
If those allegations are proven, the legal question begins before model training does.
Training and obtaining the training copy are not necessarily the same legal act
Modern language models generally do not operate as conventional searchable archives of every source file used during training. The training process converts patterns in those datasets into model parameters.
That technical transformation does not automatically resolve how the source files were obtained.
AI developers benefit from treating acquisition and training as parts of one transformative pipeline. Copyright owners have strong incentives to separate them into distinct acts with distinct legal consequences.
The complaint also raises the memorization problem
The publishers say Claude has reproduced protected lyrics verbatim or nearly verbatim in some responses.
That is a separate issue from dataset acquisition. It concerns memorization: the possibility that a sufficiently large model retains sequences closely enough to reproduce substantial portions of training material when prompted in particular ways.
Labs already use filtering and training techniques intended to reduce that behavior. Copyright litigation gives them another reason to care about how reliably those systems work.
Copyright metadata creates another layer of exposure
The plaintiffs also allege that identifying copyright-management information was removed from some works.
They seek damages that can reach $25,000 for certain alleged violations involving that information, alongside statutory copyright damages of up to $150,000 per willfully infringed work.
Those figures are requested statutory limits, not a judgment. Anthropic has not been found liable.
With tens of thousands of compositions allegedly involved, however, even a fraction of the maximum calculation becomes material.
Fair use is becoming the defining legal infrastructure question for AI
Several major US cases are now testing whether training frontier models on copyrighted works without a license can qualify as fair use.
The US government has supported AI companies' fair-use arguments in parts of that wider debate.
That position does not necessarily settle a case in which plaintiffs allege the underlying copies were obtained unlawfully. A court could view lawful acquisition and transformative training as separate questions.
Licensed AI music provides the industry's counterexample
The music business is simultaneously showing what a negotiated alternative looks like. Warner Music has moved from litigation against Suno toward licensed collaboration on new generative-music models.
Those arrangements are designed around authorized catalogs, artist participation and new forms of compensation.
Sony Music and other industry groups are also supporting labeling frameworks intended to distinguish human-led recordings, AI-assisted works and fully synthetic material.
The industry's position is becoming more specific than simply opposing artificial intelligence: permission and provenance are turning into the central requirements.
Anthropic now has to defend the data pipeline as well as the model
Anthropic says it disagrees with the publishers' claims and intends to defend itself vigorously.
This case alone is unlikely to settle every question surrounding copyrighted training data. Too many parallel lawsuits are moving through US courts.
It could still establish an important boundary. Even if some forms of AI training eventually qualify as fair use, that does not automatically mean every method used to obtain the training material will receive the same protection.