Sony and Warner sue Anthropic over tens of thousands of songs

Copyright litigation against AI moves from books to music, with a new charge that no longer disputes where the material came from but what was stripped from it before training.

Generated automatically · sources linked · no prior human review

Copyright litigation against artificial intelligence has left books behind and moved into music. On Friday night, Sony Music Publishing and Warner Chappell sued Anthropic in the U.S. District Court for the Northern District of California, and they also named its co-founders, Dario Amodei and Benjamin Mann, personally. The world’s two largest music publishers describe a “brazen campaign of illegal torrenting, scraping and downloading of protected works” and identify tens of thousands of compositions: considerably more than the music industry’s previous lawsuits, which totaled just over twenty thousand works.

The lawsuit points to three ways the material was obtained: torrents from Library Genesis and Pirate Library Mirror, scraping of licensed lyrics sites such as MusixMatch and LyricFind, and use of the public corpora Common Crawl and Books3. There are four counts, and the one worth watching is the fourth: removal or alteration of copyright management information, that is, the authorship and ownership data that travels attached to each work. That count does not depend on where the corpus came from but on what was stripped from it before training, and through that route the case could end up reaching training itself. The publishers are seeking a jury trial and the statutory maximums: up to $150,000 per willfully infringed work and $25,000 for each removal. Anthropic says it disagrees and will defend itself vigorously.

For Latin America, the problem is an empty chair. The publishing rights to the bulk of reggaeton, regional Mexican music, sertanejo and cumbia are administered by Sony and Warner, not by the people who wrote the songs: it is the publishers who are litigating, in a California court, the conditions under which that repertoire can be used for training. And they come to this case three days after becoming shareholders in Stability AI. They litigate on one side and co-invest on the other; in neither move is there a Latin American party sitting at the table.

Also today

In the region

The weekend brought no moves from the region itself: no ministries, data authorities, multilateral organizations or research centers in the region published anything within the window. What did move is the place where things that affect the region are decided without its participation. The Sony and Warner lawsuit defines the rules for training on Latin American repertoire with the publishers as plaintiffs and no composers from the region in the case. The Nature Human Behaviour study supplies the argument that the regional discussion on linguistic sovereignty was missing, just as several states are starting to put writing assistants in schools and public service counters: what is lost with automatic polishing is not content but identity, and that is not fixed by expanding language coverage. And OpenAI’s cutoff of Cursor shows that access to models can end up depending on who bought whom: the way out left to a team in Bogotá or São Paulo is to use its own API key and pay per use, in dollars, for what the subscription used to spare it. On the pending agenda, nothing has changed: PL 2338/2023 in Brazil still has no vote date, and Mexico still has no general AI law while sector-specific regulation advances.

Launches

  • Gemini Notebook introduces flexible compute-based limits — starting September 2, the quota is no longer counted in daily generations but measured by the compute each task consumes, with a reset every five hours and a weekly cap. The free plan keeps the standard limit, AI Plus doubles it, Pro quadruples it and Ultra reaches five or twenty times. It matters because it is the most widely used document synthesis tool in universities and public offices in the region precisely because the free plan was enough, and a limit measured in compute penalizes intensive use of sources, which is exactly the pattern of someone doing research with many documents.
  • LAION-BVD, an open video corpus of 10 million hours — 55 million clips with automatic video and audio descriptions, plus 300 million still images, derived from 1.3 billion CommonCrawl URLs; open, but licensed for research only. Not one of its 55 million descriptions was written by a person: they were drafted by a 2-billion-parameter model, in twenty words or fewer. It is not a tool for direct use but raw material, and it allows a university group in the region to pretrain or evaluate without negotiating access with anyone. It also serves as an object of audit: the paper does not publish a breakdown of how much of the corpus is in Spanish or Portuguese.

Threads we’re following

This adds to the copyright story we have been following. The previous chapter was the $1.5 billion settlement in the Bartz case, which ended up penalizing only one thing: where the corpus had come from. That payout came to about $3,000 per work, and no Latin American author or publisher appeared in it. The count for removal of copyright management information that appears today opens a different front, because it no longer asks about the origin of the material but about what was removed from it before training. If it succeeds, the discussion stops being about how the files were obtained and becomes about training itself.


If automatic polishing preserves almost all of the meaning but erases six percentage points of the signals of who wrote a text, at what point does a state that puts writing assistants in its schools and public service counters stop offering help and start imposing a voice?

About this entry. It is generated automatically from public sources, without human review before publication. It may contain errors of interpretation or summary; please check each story against its original source (the links lead there) before citing it or making decisions based on it.

Doble Click is written with Anthropic models.

Spotted an error? Report it

Tell us what's wrong, quoting the sentence if you can and, if you have it, the source that corrects it. An automated process reviews reports every night: if the error is verified, the page is corrected and a correction note is added at the bottom.

Your email is optional: we only use it if we need more context about the report. It doesn't subscribe you to the newsletter.