Showing posts with label AI training data. Show all posts
Showing posts with label AI training data. Show all posts

Wednesday, August 5, 2026

Why is Anthropic destroying books?; The Guardian, August 5, 2026

 Kathryn James, The Guardian; Why is Anthropic destroying books?

"Should we be surprised that destroying printed texts seemed easier to Anthropic than working with their human authors?"

Legal Battle Over U.S. Copyright Chief Heats Up Again; Publishing Perspectives, August 4, 2026

Andrew Albanese, Publishing Perspectives; Legal Battle Over U.S. Copyright Chief Heats Up Again

"In the United States, a legal battle over the future of the nation’s top copyright officer, Register of Copyrights Shira Perlmutter, is heating up again. In a legal filing last week, Administration lawyers told a Washington D.C. court that a recent Supreme Court decision bolstered their case that President Trump has the power to fire Perlmutter. But in a filing of their own, lawyers for Perlmutter reiterate that the president lacks such authority, arguing that, as an appeals court ruled last fall, the law gives that authority to the U.S. Librarian of Congress.

The legal drama began last May, when Trump purportedly fired Perlmutter, just two days after the shock firing of Librarian of Congress Carla Hayden. The firing surprised and outraged stakeholders in the copyright and Intellectual Property communities, who have given Perlmutter high marks for her work at the office, including significant progress on a much-needed modernization effort.

More concerning, however, Perlmutter’s attempted firing came immediately following the prepublication release of the Copyright Office’s third and final part of a wide-ranging AI review, which argued for the rights of copyright owners—an opinion that appears to clash with the president’s AI goals."

Tuesday, August 4, 2026

Authors weigh in on $1.5 billion Anthropic AI copyright settlement; WGCU, August 3, 2026

Mike Kiniry, WGCU ; Authors weigh in on $1.5 billion Anthropic AI copyright settlement

"As Generative AI language models have entered the scene in recent years, a wave of copyright lawsuits has arisen in response, brought by authors and publishers. These lawsuits hinge on whether downloading and ingesting millions of copyrighted books without explicit permission to train Large Language Models constitutes copyright infringement or is protected as fair use.

Authors and publishers argue that it is infringement — particularly when AI developers illegally pirate or copy their books to help train their language models. AI companies argue that reading and learning from text is transformative and therefore falls under fair use.

In one class action lawsuit that was recently settled, the AI Company Anthropic agreed to pay $1.5 billion dollars in a landmark copyright infringement settlement. It's one of the biggest in U.S. history.

There are other similar high-profile cases, including one by publishing houses including Hachette, Macmillan, and McGraw Hill, along with bestselling novelist and former President of the Author's Guild Scott Turow against Meta and its CEO, Mark Zuckerberg and another against Google. Those cases are ongoing.

The Anthropic settlement means payments of roughly $3,100 to the authors and publishers of nearly half a million books, including our guests. We have a conversation about that settlement, and other pending cases, and what this all means for the publishing world.

Guests:

Marty Ambrose-McLaughlin is an award-winning author and English instructor at Florida Southwestern State College
Scott Turow is a writer and former attorney. He is the author of fourteen works of fiction, including Presumed Innocent and his most recent, Presumed Guilty which was published in 2025."

Tuesday, July 28, 2026

Authors have mixed feelings about the $1.5B Anthropic copyright infringement ruling; NPR, July 27, 2026

 , NPR; Authors have mixed feelings about the $1.5B Anthropic copyright infringement ruling

"Graeber is among the more than 300,000 writers involved in the suit who may soon be getting a modest windfall. A federal judge in San Francisco rubber stamped a $1.5 billion settlement in July resulting from a landmark class action lawsuit the authors brought against the AI company Anthropic two years ago...

AI companies often invoke the fair use doctrine – which enables the use of copyrighted works without the copyright holder's consent in some situations – as they try to make the case in court for training their models on these materials...

Chinese AI companies often use a technique to build their models called "AI distillation." This involves feeding their models the outputs generated by other AI models, often high-quality U.S.-based ones like OpenAI's GPT-4 or Anthropic's Claude, instead of directly training them on pirated copies of books by American authors...

One possible way for authors to get a fairer shake in the age of AI could be through the licensing of their work to AI companies...

There are already some such deals between publishers and AI companies in place, such as Perplexity AI's agreement with media entities like the Los Angeles Times and Le Monde to license content for the training of its models. There are also online licensing marketplaces, such as Created by Humans."


Friday, July 24, 2026

UK's Bloomsbury among beneficiaries of $1.5 billion Anthropic copyright lawsuit settlement; Reuters, July 22, 2026

Reuters ; UK's Bloomsbury among beneficiaries of $1.5 billion Anthropic copyright lawsuit settlement

"Britain's Bloomsbury ​Publishing confirmed on Wednesday it was among ‌the beneficiaries of a landmark $1.5 billion settlement that resolves claims artificial intelligence ​company Anthropic used copyrighted books ​to train its AI models without ⁠purchasing the content.

Here are some ​more details:

  • Bloomsbury said a U.S. court ​identified 14,087 of its titles covered by the settlement, with proposed compensation of ​about $3,000 per title, split equally ​between the author and publisher...
  • The settlement ​is the largest known copyright payout ​in ⁠U.S. history."

Wednesday, July 22, 2026

Saturday, July 18, 2026

Hackers Expose How AI Music App Suno Stole Decades Worth of Copyrighted Music; Futurism, July 17, 2026

, Futurism; Hackers Expose How AI Music App Suno Stole Decades Worth of Copyrighted Music

The evidence is damning.

"A hack revealed in detail how AI music generating app Suno scraped millions of songs, likely including copyrighted ones, from across the web to feed into its AI model, 404 Media reports.

Suno, which is currently embroiled in multiple ongoing copyright lawsuits, has already admitted in response to legal action that it used “essentially all music files of reasonable quality that are accessible on the open internet” to train its music-generating AI."


Friday, July 10, 2026

The Work of Helping A.I. Destroy Work; The New York Times, July 10, 2026

  , The New York Times; The Work of Helping A.I. Destroy Work

"Every day, Mercor, a start-up that sells training data to artificial intelligence companies, pays 30,000 contractors more than $4 million to help make their jobs, and those of their colleagues, obsolete.

It’s gig work, but for professionals with rarefied skills. One recent Mercor posting offered $225 an hour for a voice actor able to maintain a customer service persona in fluent Hebrew. Another sought a Ph.D. physicist with a specialization in general relativity, astrophysics or cosmology. A third listing wanted a physician with more than three years of experience in the Rwandan primary care medical system.

Mercor and a handful of similar start-ups are the primary middlemen in a supply chain of “human data” that may power the next generation of A.I. As OpenAI, Anthropic and other major ventures compete to become the industry’s dominant platform, the market for premium data that has been vetted by experts is exploding.

No longer do the A.I. companies need armies of low-paid workers, often overseas, to do rote tasks like tag images of cars or transcribe audio. They need mathematicians to annotate proofs, lawyers to mark up briefs and professors to grade essays. That’s what Mercor and its rivals supply. To use the parlance of the industry, data labeling has moved up the “value chain,” and the start-ups that offer this service have become some of the fastest growing in Silicon Valley."

OpenAI may have made a fatal misstep in copyright fight with news orgs; Ars Technica, July 9, 2026

ASHLEY BELANGER  , Ars Technica; OpenAI may have made a fatal misstep in copyright fight with news orgs

"OpenAI is facing calls for “serious sanctions” after fighting to keep news organizations from snooping through millions of logs to find evidence of users skirting their paywalls by prompting ChatGPT to regurgitate their articles.

This evidence is considered among the most important to both sides, potentially either dooming OpenAI as an infringer or exonerating its chatbot technology as a transformative fair use of news sites’ content."

Thursday, July 2, 2026

Microsoft Shareholder Sues Top Brass for AI Copyright Claims; Bloomberg Law, July 1, 2026

 , Bloomberg Law; Microsoft Shareholder Sues Top Brass for AI Copyright Claims

"Microsoft Corp.'s executives and board directors misled shareholders in statements concealing its artificial intelligence tools were trained on copyrighted material, a new investor lawsuit said. 

The tech giant’s false statements about its AI strategy and violations of intellectual property law caused substantial damage to Microsoft and its shareholders, according to Eric Anderson’s stockholder derivative lawsuit filed Tuesday in the US District Court for the Western District of Washington."

Tuesday, June 30, 2026

Ford rehires human engineers after AI fails to match quality checks; BBC, June 29, 2026

 Liv McMahon , BBC; Ford rehires human engineers after AI fails to match quality checks

"Ford says it has hired back some human engineers after AI failed to match their skills and experience.

In a bid to reap the benefits of the tech, which developers claim can cut costs and boost productivity, the US carmaker adopted it across some parts of its operations including for quality checks.

But, according to Bloomberg, its executives said the firm has rehired more than 300 "veteran" quality inspectors in recent years to make up for the pitfalls of automated systems.

"Artificial intelligence is a fantastic tool, but it's only as good as the information you use to train it," Charles Poon, vice president of vehicle hardware engineering, told reporters.

"Over prior years, we didn't pay as much attention as we should have to the experience of our most knowledgeable engineers that have been with us through many product cycles," he said.

The US automaker is among many to have seized on the buzz around AI, particularly amid Wall Street fervour about the tech's potential to increase margins."

Monday, June 29, 2026

NYT slams Microsoft for building copyright-infringing supercomputer for OpenAI; Ars Technica, June 26, 2026

ASHLEY BELANGER , Ars Technica; NYT slams Microsoft for building copyright-infringing supercomputer for OpenAI

"In a heavily redacted court filing Thursday, The New York Times proposed to amend its copyright complaint against OpenAI and Microsoft to clarify a claim and allege that Microsoft actively encouraged OpenAI to steal NYT works by building a bespoke supercomputing system ranked among the most powerful in the world."

Saturday, June 27, 2026

How Teaching A.I. to Speak Cajun Can Help Save a Language; The New York Times, June 27, 2026

, The New York Times; How Teaching A.I. to Speak Cajun Can Help Save a Language

By feeding centuries-old nursery rhymes and folklore recordings into their own model, linguists in Louisiana hope to help a community control its digital destiny.

"Louisiana French, the oral dialect of which Balfa was a cultural guardian, is part of the Bayou’s societal DNA, a link to its history, music and identity. Today, Caffery described the language as struggling and endangered, a notion reinforced by Alexa’s overlooking Balfa.

In response, Caffery assembled a small team at the center to train its own language learning model in automatic speech recognition for Louisiana French, drawing from a trove of historical artifacts and interviews.

Over the months, as the learning language model is trained on bits of the language — such as an old-age French nursery rhyme — it brings centuries-old dialect closer into the digital age."

Thursday, June 25, 2026

The New York Times Amends Lawsuit Against OpenAI and Microsoft; The New York Times, June 25, 2026

 , The New York Times; The New York Times Amends Lawsuit Against OpenAI and Microsoft

"The New York Times amended its lawsuit against OpenAI and Microsoft on Thursday, modifying one claim against Microsoft and dropping another against OpenAI, according to a legal filing in federal court...

In a filing in the U.S. District Court for the Southern District of New York on Thursday, The Times accused Microsoft of encouraging OpenAI to train its A.I. systems using copyrighted articles from The Times and of providing services designed to help with this training.

The Times also dropped a claim from its original lawsuit, filed in 2023, accusing OpenAI of “secondarily” infringing on its copyrights because it did not prevent consumers and businesses from generating copyrighted material using A.I."

Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft; New Jersey Globe, June 24, 2025

David Wildstein, New Jersey Globe ; Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft

"The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including ChatGPT and Microsoft Copilot, without permission or compensation.

The publishers, represented by Platkin LLP, a law firm founded earlier this year by former New Jersey Attorney General Matthew J. Platkin, contend that OpenAI and Microsoft unlawfully appropriated original news content to build their AI systems, violating the Copyright Act and threatening the future of local journalism.

The lawsuit also alleges that OpenAI knowingly stripped copyright management information from publishers’ work — including author bylines, copyright notices, and terms of use information — in violation of the Digital Millennium Copyright Act.

The complaint cites remarks by OpenAI founder Sam Altman, who acknowledged during testimony before the British House of Lords that it would be “impossible to train today’s leading AI models without using copyrighted materials.”"

Tuesday, June 23, 2026

Archiving with AI; Library Journal, June 8, 2026

 Matt Enis, Library Journal; Archiving with AI

"AI companies are offering some libraries funding for digitization projects, but archives and special collections are working through how to manage projects responsibly

“Imagine a world where you know things but cannot say where you learned them,” begins “Memory Without Origin,” a paper published in April by University of Virginia (UVA) Dean of Libraries and University Librarian Leo S. Lo. This isn’t a hypothetical question, Lo notes, it’s a predictable consequence if libraries allow generative artificial intelligence (AI) to ingest archival materials as training data without requiring provenance conditions. And libraries, which could always use funding for projects involving digitization, special collections, and archives, are being approached by AI companies with deep pockets.

“They’ve been approaching a lot of larger research libraries, including Oxford and many more,” Lo tells LJ. (Oxford’s Bodleian Libraries began a digitization pilot project funded by ChatGPT maker OpenAI last year.) “Usually the offer is: they will pay you to digitize materials—which we want, because we want to make them more accessible—and in return, depending on the deal…they would like to have the data to train their AI models.”

These partnerships can benefit both parties, but for libraries, the consequences of getting these arrangements wrong “are more permanent than anything the profession has previously encountered,” Lo writes. “Once archival materials are absorbed into foundation model weights, no subsequent institutional action can remove them from the model.” If proper care isn’t taken, that information becomes unmoored from its former context within an archive."

Friday, June 19, 2026

Millions of Copyrighted Songs Were Fed to AI Music Generators – Now There’s Proof; Gadget Review, June 16, 2026

Al Landes , Gadget Review; Millions of Copyrighted Songs Were Fed to AI Music Generators – Now There’s Proof

Atlantic databases name 21 million tracks fed to Suno and rivals as Sony, UMG, and Warner seek $150,000 per song in damages

"Searchable databases verify roughly 21 million copyrighted songs trained AI music generators.

Sony, UMG, and Warner lawsuits seek up to $150,000 per song from Suno and Udio.

HarmonyCloak tool lets artists protect songs by adding inaudible AI-blocking audio perturbations.

Millions of copyrighted songs — including chart-topping hits — verifiably trained AI music generators, and now there are searchable databases to prove it. The Atlantic, through an investigation by staff writer Alex Reisner, published four catalogs documenting exactly which music fed these models:"

Tuesday, June 16, 2026

Publishers Sue WeLib for Copyright Infringement; Publishers Weekly, June 16, 2026

 Jim Milliot , Publishers Weekly; Publishers Sue WeLib for Copyright Infringement

"Fresh off of last month’s victory against pirate web site Anna’s Archive, 13 publishers across all segments of the industry have allied to sue yet another pirate site, WeLib, for copyright infringement.

The suit, filed in the U.S. District Court for the Southern District of New York, charges that the operators of WeLib “ copied the source code and most of the contents of” Anna's Archive."

The plaintiffs include the Big Five, Cengage, Elsevier, McGraw Hill, Pearson, Taylor & Francis, and Wiley.

“Defendants boast that they have reproduced ‘an endless collection of literature, research papers, and education materials,’ none of which they own or have licensed,” the complaint alleges. 

According to its website and repeated in the lawsuit, WeLib hosts over 43 million books and 98 million papers, and its stolen collection of literary works has purportedly attracted over 80,000 active monthly users. According to the website, WeLib’s users have illegally accessed over 51 million books in the last month alone, or an average of over 1.7 million books per day."