Why AI companies are purchasing old books, scanning them, and destroying the evidence
In a bizarre twist for the digital-first artificial intelligence industry, developers are turning to physical destruction to fuel their virtual brains. A report published by Novara Media on July 29, 2026, reveals that AI firms are quietly buying up physical copies of out-of-print and copyrighted books, scanning them to harvest pristine AI training data, and then physically destroying the volumes. This aggressive tactic represents a desperate bid to bypass digital rights management (DRM) and paywalls as the industry face-plants into a severe data shortage.
The Loophole in the Paper Trail
AI developers are currently facing an existential data wall. With internet-scraping lawsuits piling up and high-quality web data exhausted, firms are seeking pristine, non-synthetic sources of human language. Purchasing physical books from secondary markets, slicing their spines to feed them into high-speed scanners, and shredding the evidence is a legally grey but highly effective workaround. Under the traditional "first sale doctrine," a buyer is legally permitted to sell, donate, or destroy their physical copy of a book—even if reproducing its digital text for training commercial models remains a major legal battleground.
Escaping the Synthetic Feedback Loop
Destructive scanning—often called "book slicing"—is the fastest way to get perfectly flat, high-resolution pages for optical character recognition (OCR) software. But the destruction of the physical books also serves another purpose. By eliminating the physical evidence, companies make it far harder for publishers to audit what texts have been digitized. More importantly, digitizing rare or physical-only texts gives these firms exclusive access to a pool of "clean" human language untouched by the poisoning effects of synthetic, AI-generated content currently flooding the web.
"They aren't just digitizing culture; they are consuming it. This is a literal enclosure of the physical commons to train corporate algorithms."
Novara Media, July 29, 2026
The Cultural Preservation Backlash
While the story has seen low initial engagement on networks like Hacker News, the implications for librarians, archivists, and preservationists are massive. By systematically purchasing and destroying obscure, historical, or out-of-print literature, tech companies are removing physical copies from circulation entirely. It represents a modern-day digital enclosure movement, where physical knowledge is privatized, digested, and then destroyed to feed proprietary models.
If the future of machine intelligence depends on the literal destruction of our physical history, we have reached a bizarre inflection point. AI companies are no longer just indexing the world's information—they are burning the library to build the model.
This article was ultrathought.
Get breaking news, funding rounds, and analysis delivered to your inbox. Free forever.