, The Telegraph; Cultural barbarism: How AI companies are destroying the world’s books
In its attempt to boost AI’s power, Silicon Valley is buying millions of rare editions, scanning them and then shredding the originals
My Bloomsbury book "Ethics, Information, and Technology" was published on Nov. 13, 2025. Purchases can be made via Amazon and this Bloomsbury webpage: https://www.bloomsbury.com/us/ethics-information-and-technology-9781440856662/
Chris Stokel-Walker , The Telegraph; Cultural barbarism: How AI companies are destroying the world’s books
In its attempt to boost AI’s power, Silicon Valley is buying millions of rare editions, scanning them and then shredding the originals
Matt Enis, Library Journal; Archiving with AI
"AI companies are offering some libraries funding for digitization projects, but archives and special collections are working through how to manage projects responsibly
“Imagine a world where you know things but cannot say where you learned them,” begins “Memory Without Origin,” a paper published in April by University of Virginia (UVA) Dean of Libraries and University Librarian Leo S. Lo. This isn’t a hypothetical question, Lo notes, it’s a predictable consequence if libraries allow generative artificial intelligence (AI) to ingest archival materials as training data without requiring provenance conditions. And libraries, which could always use funding for projects involving digitization, special collections, and archives, are being approached by AI companies with deep pockets.
“They’ve been approaching a lot of larger research libraries, including Oxford and many more,” Lo tells LJ. (Oxford’s Bodleian Libraries began a digitization pilot project funded by ChatGPT maker OpenAI last year.) “Usually the offer is: they will pay you to digitize materials—which we want, because we want to make them more accessible—and in return, depending on the deal…they would like to have the data to train their AI models.”
These partnerships can benefit both parties, but for libraries, the consequences of getting these arrangements wrong “are more permanent than anything the profession has previously encountered,” Lo writes. “Once archival materials are absorbed into foundation model weights, no subsequent institutional action can remove them from the model.” If proper care isn’t taken, that information becomes unmoored from its former context within an archive."
Donna Ferguson, The Guardian; A bonanza for fans of the natural world: the digital library sharing 64m pages of scientific knowledge with everyone
"Over the past 20 years, more than 64m pages have been made freely available through the Biodiversity Heritage Library (BHL) – a digital treasure trove for fans of the natural world. More than 680 museums, universities, libraries and scientific institutions from China, Singapore, Australia and New Zealand to Europe, Africa, Mexico, Canada and the US, have contributed to the library.
This week, a report from Royal Botanic Gardens (RBG), Kew revealed the crucial role digitisation is playing in “transforming our ability to understand and respond to the climate and biodiversity crises”, but it was the creation of the BHL 20 years ago that first demonstrated how bringing centuries of scientific knowledge online can unlock transformative discoveries and insights about the natural world.
David Iggulden, who chairs the BHL executive committee alongside his job as head of data and digital, library and archives at RBG Kew, describes the library as an invaluable and “absolutely essential” resource for scientists in the field. But it is also used by scientific researchers, environmental historians, educators, art historians, artists, citizen scientists and members of the public who – like Iggulden – simply enjoy browsing its contents on a rainy weekend.
“I just get caught up in it sometimes, looking at the various collections,” he says. “I think it’s amazing that we can explore such a vast array of different collections from very different institutions.”
As well as published biodiversity literature and journals, there are letters, illustrations, climate records, field diaries, ecosystem profiles, distribution records and manuscripts containing the original collecting stories of a particular species or detailing voyages of discovery."
Rory Carroll , The Guardian; Lost copy of seventh-century poem in Old English discovered at Rome library
"“This discovery is a testament to the power of libraries to facilitate new research by digitising their collections and making them freely available online,” she said.
Andrea Cappa, head of manuscripts and rare books at the Rome library, said the institution was digitising holdings from Italy’s National Centre for the Study of the Manuscript, which will give researchers access to more than 40m images.
Riccardo Fangarezzi, head of archives at the abbey in Nonantola, said he looked forward to further discoveries. “The present times may be rather dark, yet such intellectual contributions are genuine rays of sunlight: the continent is less isolated,” he said.
The poet Paul Muldoon translated Caedmon’s Hymn into contemporary English in a 2016 anthology of British poetry. The opening lines read:
“Now we must praise to the skies, the Keeper of the heavenly kingdom,
The might of the Measurer, all he has in mind,
The work of the Father of Glory, of all manner of marvel.”"
Chloe Veltman, NPR ; Boston Public Library aims to increase access to a vast historic archive using AI
"Boston Public Library, one of the oldest and largest public library systems in the country, is launching a project this summer with OpenAI and Harvard Law School to make its trove of historically significant government documents more accessible to the public.
The documents date back to the early 1800s and include oral histories, congressional reports and surveys of different industries and communities...
Currently, members of the public who want to access these documents must show up in person. The project will enhance the metadata of each document and will enable users to search and cross-reference entire texts from anywhere in the world.
Chapel said Boston Public Library plans to digitize 5,000 documents by the end of the year, and if all goes well, grow the project from there...
Harvard University said it could help. Researchers at the Harvard Law School Library's Institutional Data Initiative are working with libraries, museums and archives on a number of fronts, including training new AI models to help libraries enhance the searchability of their collections.
AI companies help fund these efforts, and in return get to train their large language models on high-quality materials that are out of copyright and therefore less likely to lead to lawsuits. (Microsoft and OpenAI are among the many AI players targeted by recent copyright infringement lawsuits, in which plaintiffs such as authors claim the companies stole their works without permission.)"
Michael Hiltzik , Los Angeles Times; Column: A Faulkner classic and Popeye enter the public domain while copyright only gets more confusing
"The annual flow of copyrighted works into the public domain underscores how the progressive lengthening of copyright protection is counter to the public interest—indeed, to the interests of creative artists. The initial U.S. copyright act, passed in 1790, provided for a term of 28 years including a 14-year renewal. In 1909, that was extended to 56 years including a 28-year renewal.
In 1976, the term was changed to the creator’s life plus 50 years. In 1998, Congress passed the Copyright Term Extension Act, which is known as the Sonny Bono Act after its chief promoter on Capitol Hill. That law extended the basic term to life plus 70 years; works for hire (in which a third party owns the rights to a creative work), pseudonymous and anonymous works were protected for 95 years from first publication or 120 years from creation, whichever is shorter.
Along the way, Congress extended copyright protection from written works to movies, recordings, performances and ultimately to almost all works, both published and unpublished.
Once a work enters the public domain, Jenkins observes, “community theaters can screen the films. Youth orchestras can perform the music publicly, without paying licensing fees. Online repositories such as the Internet Archive, HathiTrust, Google Books and the New York Public Library can make works fully available online. This helps enable both access to and preservation of cultural materials that might otherwise be lost to history.”"
JON BLISTEIN, Rolling Stone; KATHLEEN HANNA, TEGAN AND SARA, MORE BACK INTERNET ARCHIVE IN $621 MILLION COPYRIGHT FIGHT
"Kathleen Hanna, Tegan and Sara, and Amanda Palmer are among the 300-plus musicians who have signed an open letter supporting the Internet Archive as it faces a $621 million copyright infringement lawsuit over its efforts to preserve 78 rpm records...
The lawsuit was brought last year by several major music rights holders, led by Universal Music Group and Sony Music. They claimed the Internet Archive’s Great 78 Project — an unprecedented effort to digitize hundreds of thousands of obsolete shellac discs produced between the 1890s and early 1950s — constituted the “wholesale theft of generations of music,” with “preservation and research” used as a “smokescreen.” (The Archive has denied the claims.)
While more than 400,000 recordings have been digitized and made available to listen to on the Great 78 Project, the lawsuit focuses on about 4,000, most by recognizable legacy acts like Billie Holiday, Frank Sinatra, Elvis Presley, and Ella Fitzgerald. With the maximum penalty for statutory damages at $150,000 per infringing incident, the lawsuit has a potential price tag of over $621 million. A broad enough judgement could end the Internet Archive.
Supporters of the suit — including the estates of many of the legacy artists whose recordings are involved — claim the Archive is doing nothing more than reproducing and distributing copyrighted works, making it a clear-cut case of infringement. The Archive, meanwhile, has always billed itself as a research library (albeit a digital one), and its supporters see the suit (as well as a similar one brought by book publishers) as an attack on preservation efforts, as well as public access to the cultural record."
U.S. Copyright Office; Copyright Office Launches Digitized Copyright Historical Record Books Collection
"The Copyright Office today launched the first release of the digitized Copyright Historical Record Books Collection. “The Copyright Office holds the world’s most comprehensive collection of records of copyright ownership,” said Register of Copyrights Shira Perlmutter. “Today’s release of the first batch of our digitized historical record books will ensure that these records are preserved for future research and that anyone can access them from anywhere.”
This collection is a preview of digitized versions of historical record books that the Office plans to incorporate into its Copyright Public Record System (CPRS), currently in public pilot. The collection will eventually include images of copyright applications and other records bound in books dating from 1870 to 1977. This first release includes 500 record books containing registration applications for books from 1969 to 1977, with a majority of the record books being the most recent volumes from 1975 to 1977. The collection is being digitized using the Copyright Office’s internal administrative classification system in reverse chronological order. There will be periodic updates as record books are digitized and added to the collection."
"The age of digitisation has opened new doors to distribution of information including for libraries and archives. However, librarians and archivists are often confronted with risk of liability for copyright infringement, nationally and in cross-border activities. This week, they asked the World Intellectual Property Organization copyright committee to provide them not only with some exceptions to copyright, but with protection against liability. The WIPO Standing Committee on Copyright and Related Rights (SCCR) is taking place from 14-18 November. On the SCCR agenda is copyright exceptions and limitations for libraries and archives. On 17 November, librarians and archivists took the floor to explain why an international standard protecting them against liability is indispensable."
'Decorating with Christmas tree candles was in its commercial prime for about 50 years — from roughly 1870, gradually tapering off through the 1920s and 30s, although the U.S. Patent Office was still awarding patents for Christmas tree candle holders as late as 1945. Studying patent drawings and reading the filings from the U.S. Patent Office is a good way to trace the design developments and get a feel for how the Christmas tree candle holders evolved through the years before they were overtaken by electric Christmas tree lights. The following selection is excerpted from our online Gallery of Christmas Tree Candle Holder Patents, which is the largest collection available on the web. Check out the gallery and see varieties of Christmas tree candle holders, clips and pendulums that are available to buy online."
"Harvard Library recently released a comprehensive literature review on orphan works and copyright in “an attempt to solve the legal complexities of the orphan works problem by identifying no-risk or low-risk ways to digitize and distribute orphan works under U.S. copyright law. The project’s goal is to help clear the way for U.S. universities, libraries, archives, museums, and other cultural institutions to digitize their orphan works and make the digital copies open access.” The review, “Digitizing Orphan Works: Legal Strategies to Reduce Risks for Open Access to Copyrighted Orphan Works,” was written by David Hansen, clinical assistant professor of law and faculty research librarian at the UNC (University of North Carolina) School of Law. Its goal is to “change the face of the orphan-works problem in the United States.”"
"West Virginia University Press and the WVU Libraries have launched West Virginia History: An Open Access Reader, a free, online collection of previously published essays drawn from the journal West Virginia History and other WVU Press publications. The collection covers the history of the territory that became West Virginia from European settlement to mountaintop removal, and is especially suitable for use in courses on state history. It is available at https://textbooks.lib.wvu.edu/index.html. “I love everything about this project – serving the needs of students, sharing the history of West Virginia, harnessing the power of technology, and collaborating between West Virginia University and Marshall,” WVU President Gordon Gee said. “This is the kind of responsive and innovative work that we want to become the ‘new normal’ for higher education in meeting the needs of our state.”"
"Hundreds of stunning images from black history, drawn from old negatives, have long been buried in the musty envelopes and crowded bins of the New York Times archives. None of them were published by The Times until now. Were the photos — or the people in them — not deemed newsworthy enough? Did the images not arrive in time for publication? Were they pushed aside by words here at an institution long known as the Gray Lady?... Every day during Black History Month, we will publish at least one of these photographs online, illuminating stories that were never told in our pages and others that have been mostly forgotten... Many of these photographs, and their stories, are equally intriguing. But the collection is far from comprehensive. There are gaps, for many reasons."
"But the game is what you might call a marketing teaser for a major redistribution of property, digitally speaking: the release of more than 180,000 photographs, postcards, maps and other public-domain items from the library’s special collections in downloadable high-resolution files — along with an invitation to users to grab them and do with them whatever they please. Digitization has been all the rage over the past decade, as libraries, museums and other institutions have scanned millions of items and posted them online. But the library’s initiative (nypl.org/publicdomain), which goes live on Wednesday, goes beyond the practical questions of how and what to digitize to the deeper one of what happens next... A growing number of institutions have been rallying under the banner of “open content.” While the library’s new initiative represents one of the largest releases of visually rich material since the Rijksmuseum in Amsterdam began making more than 200,000 works available in high-quality scans free of charge in 2012, it’s notable for more than its size. “It’s not just a data dump,” said Dan Cohen, the executive director of the Digital Public Library of America, a consortium that offers one-stop access to digitized holdings from more than 1,300 institutions. The New York Public has “really been thinking about how they can get others to use this material,” Mr. Cohen continued. “It’s a next step that I would like to see more institutions take.” Most items in the public-domain release have already been visible at the library’s digital collections portal. The difference is that the highest-quality files will now be available for free and immediate download, along with the programming interfaces, known as APIs, that allow developers to use them more easily."