← Back to Main Page

AI Companies are Destroying Books and Why the Library of Congress Can't Save Them

by Jeremy C.

Money · July 29, 2026

Library of Congress employees restoring a book.

Library of Congress employees restoring a book.

There is a machine called a hydraulic cutter and it works by shearing the spine off a bound book in a single pass. It leaves a stack of loose pages that an industrial scanner then copies digitally in seconds. The paper, along with the other sheared off pieces of the book, gets destroyed. But the digital form that came from the process only exists for perhaps a few minutes. The words of the book that detailed a historical moment or perhaps some scientific insight live just long enough for a computer system built on circuits and mathematical algorithms to process and assimilate it within its memory. The exact work is purged with only the impression that work made the only thing that's left. In most respects, the book doesn't exist anymore. It's become one algorithmic impression among many.

The process to turn millions of books to algorithmic figments in this fashion was given a name. It was called Project Panama and run by the Artificial Intelligence company Anthropic, which created and runs the Claude chatbot. The aim was to bring in massive amounts of information to use as so-called training data, which increases the mathematical connections that the AI system can make across subjects. But the company was careful to couch it as training data and destroy the actual materials once training was completed since many of the works it was processing were protected by copyright. Federal courts have viewed, in some circumstances, a transformative use of purchased physical works to be acceptable and not a violation of copyright law. So in went the physical books with only the digital connections made from its processing remaining.

From Google to Anthropic

The man hired to destroy the books for Anthropic helped build Google Books. Anthropic recruited Tom Turvey, formerly head of partnerships for Google Books, in February 2024 and tasked him with obtaining "all the books in the world."

Of course, there are a number of questionable aspects to this scan and destroy strategy. But for this article, the main focus is on its potential to destroy the only remaining copies of a book or article. Many people within the book business are raising this concern the loudest. As reported in the news, a number of booksellers spoke to the Washington Post on condition of anonymity saying that rare, out-of-print titles with few surviving copies were among the stock being bought and processed for AI training. While we don't know for certain if any of the books processed so far were the last of their kind--and therefore wiped completely from existence--we do know the risk is real and is only getting worse as these types of operations continue. But the interesting thing is that the United States once had a failsafe to ensure that published works would survive no matter what happened to the publisher or the commercially distributed copies.

The founders of the United States saw the protection of written works as so important that it was embedded into the Constitution even before the Bill of Rights were. In Article I, the Constitution identified the power of Congress to promote the progress of science and useful arts by securing a limited time of exclusive right to their writings and discoveries. This is what we call the copyright system where someone creates a work and can have the exclusive right to sell it. This creates a productive incentive since the creator can profit from a work without needing to worry about others just taking it over. But of course, it's limited since that protection comes via the public's resources of courts and judges. Once that limited time is up, the country benefits by having a work that is now open for use by anyone who would like to use it. It's a delicate balance that protects property rights but also prevents permanent exclusive control of knowledge and science.

But for the system to fulfill the public end of the bargain, those written works and publications had to remain available after that time of limited control by the creator was up. And for the first one hundred years of America's existence, the country had a place where these works were delivered, housed, and indeed kept available. The delivery to the government of one copy of every copyrighted work was part of the deal to receive court-enforced protection. Initially, publishers sent their works to district courts with the State Department being the final destination for preservation. Later, it was changed to the Library of Congress. But no matter what federal entity was made responsible for receiving the works, the ultimate objective was always the same: preservation. And if that system would have continued, we wouldn't be at risk of knowledge vanishing simply because an AI company happened to use the last remaining copy for algorithmic training data. But the Library of Congress is no longer the failsafe it was meant to be.

Inside Man?

Herbert Putnam, who was the eighth Librarian of Congress and successfully pushed for the power of the Library of Congress to destroy published works in 1909, was the son of the founder of G.P. Putnam's Sons, a large publishing company. He was also the brother of the publishing industry's chief copyright lobbyist.

In 1909, Congress passed "An Act To amend and consolidate the Acts respecting copyright." Instead of the simple and direct mission of preserving all works, it allowed the librarian of Congress the discretion to choose what got preserved and what could be disposed of. There's alot more to the story about this initial blow to the protection of published works, to include the then congressional librarian's connections to a publishing giant at the time, but that will have to wait for a different article. The next and more destructive blow came with the 1976 general revision of copyright law that Congress passed and which took effect in 1978. The requirement to send a copy of each published work was watered down in some ways and completely abolished in others. Publishers and individuals could now get the protection of federal courts without necessarily needing to send a copy to be kept by public institutions. With that, the bargain that the founders struck became one-sided where the public end was often unable to be exercised. If you're promised a right to use a work after a certain date but that work was destroyed--or not even submitted in the first place--then how will you be able to exercise that right?

The arguments for allowing the librarian of Congress to destroy previously protected works varied. From the lack of space to store them to certain items being viewed as low value, arguments were made that allowed destruction to begin in 1909. Then in 1976, arguments were made to relax the requirements to send the items in the first place. From the cost of sending materials by mail to the potential for publishers losing protections simply because a copy didn't make it in the mail on time, arguments were made that allowed nothing to be sent at all in some instances. But all of those arguments and changes were built around catering to specific interests. The Library of Congress saved space and corporate publishers saved money and legal costs, but now we're all at risk of losing some pockets of knowledge forever.

Quin and Frank at the Arcade
A comic strip with nostalgia.
Frank Reading the Paper
You don't have to search the newspapers...
Gromlee Eating Pizza
...to find humor at GROMLEE.COM