Sony Music Publishing, Warner Chappell, and numerous other music publishers filed a lawsuit against Anthropic and its co-founders on Friday, accusing the AI lab of running a brazen campaign of illegally torrenting, scraping, and downloading copyrighted works to train its Claude language model.
The lawsuit, filed in the U.S. District Court for the Northern District of California, targets Anthropic CEO Dario Amodei and co-founder Benjamin Mann personally alongside the company. It represents the latest and most expansive legal challenge to how AI labs obtain the training data that powers their systems.
The publishers accuse Anthropic of blatant theft by using thousands of copyrighted musical works, including lyrics and sheet music, to develop Claude. The complaint describes illegal torrenting to obtain millions of copies of books and other works that contain protected musical content, arguing that the scale and method of acquisition went far beyond any reasonable fair use defense.
Building on Earlier Cases
The lawsuit does not exist in isolation. Some of the same legal teams behind this filing also represent Concord Music Group and Universal Music Group in a case filed in January that accused Anthropic of flagrant piracy involving 20,000 works. Those earlier suits laid the groundwork for the broader allegations now being pressed by Sony and Warner.
Perhaps most significantly, this latest action builds on the landmark Bartz v. Anthropic case, in which a group of authors accused Anthropic of using copyrighted books to train Claude. A judge in that case ruled that while it was legal for the AI lab to use copyrighted works for training, it was not legal to obtain those works through pirated sources. Anthropic was ordered to pay $1.5 billion in the largest copyright settlement in history, a ruling that sent shockwaves through the AI industry.
The Sony and Warner complaint explicitly references that decision, arguing that the same pirating methodology used in the Bartz case was applied across the music industry’s catalog. The publishers contend that Anthropic’s conduct represents not an isolated error but a systematic approach to data acquisition that prioritized speed and scale over legal compliance.
Anthropic Responds
Anthropic pushed back against the allegations in a brief statement. We disagree with the publishers’ claims and we intend to defend ourselves robustly in court, an Anthropic spokesperson wrote in an emailed statement.
The company has maintained throughout its various legal battles that its use of training data falls within established copyright principles, though the Bartz ruling complicated that argument by distinguishing between lawful use and unlawful acquisition methods. The distinction matters enormously for the AI industry, which relies on vast datasets scraped from the internet to build its models.
The Stakes for AI Training
The music publishers’ lawsuit carries implications that extend well beyond the relationship between Anthropic and the recording industry. If the court finds that Anthropic’s torrenting of copyrighted materials constitutes infringement, it could establish precedent affecting how every major AI lab sources its training data.
The case also raises questions about personal liability for AI executives. By naming Amodei and Mann as individual defendants, the publishers are arguing that the decision to use pirated data was not merely a corporate oversight but a conscious choice made by the company’s leadership. That framing, if accepted by the court, could deter other AI labs from taking similar risks with their data acquisition practices.
The music industry has been particularly aggressive in pursuing AI-related copyright claims, driven by the existential threat that AI-generated music poses to songwriters, performers, and publishers. Major labels have watched with alarm as AI tools produce increasingly convincing樑仿 of human artists, and they view the training data question as a critical front in protecting the economic value of their catalogs.
What Comes Next
The case is expected to proceed through the Northern District of California, a jurisdiction that has become the de facto courtroom for AI copyright disputes. Legal experts expect discovery to focus on the specific methods Anthropic used to obtain training data, including whether it employed torrenting networks, scraped websites hosting copyrighted content, or accessed datasets compiled by third parties from pirated sources.
With multiple parallel lawsuits now targeting the same company over overlapping allegations, Anthropic faces the prospect of litigating its data practices on several fronts simultaneously. The outcomes will shape not only the company’s future but the broader legal framework governing how AI systems are built and what data they are permitted to learn from.
discussion