BREAKING July 31, 2026 3 min read

Why AI companies are purchasing old books, scanning them, and destroying the evidence

ultrathink.ai
Thumbnail for: AI Firms Buy and Destroy Physical Books for Training Data

In a bizarre twist for the digital-first artificial intelligence industry, developers are turning to physical destruction to fuel their virtual brains. A report published by Novara Media on July 29, 2026, reveals that AI firms are quietly buying up physical copies of out-of-print and copyrighted books, scanning them to harvest pristine AI training data, and then physically destroying the volumes. This aggressive tactic represents a desperate bid to bypass digital rights management (DRM) and paywalls as the industry face-plants into a severe data shortage.

The Loophole in the Paper Trail

AI developers are currently facing an existential data wall. With internet-scraping lawsuits piling up and high-quality web data exhausted, firms are seeking pristine, non-synthetic sources of human language. Purchasing physical books from secondary markets, slicing their spines to feed them into high-speed scanners, and shredding the evidence is a legally grey but highly effective workaround. Under the traditional "first sale doctrine," a buyer is legally permitted to sell, donate, or destroy their physical copy of a book—even if reproducing its digital text for training commercial models remains a major legal battleground.

Escaping the Synthetic Feedback Loop

Destructive scanning—often called "book slicing"—is the fastest way to get perfectly flat, high-resolution pages for optical character recognition (OCR) software. But the destruction of the physical books also serves another purpose. By eliminating the physical evidence, companies make it far harder for publishers to audit what texts have been digitized. More importantly, digitizing rare or physical-only texts gives these firms exclusive access to a pool of "clean" human language untouched by the poisoning effects of synthetic, AI-generated content currently flooding the web.

"They aren't just digitizing culture; they are consuming it. This is a literal enclosure of the physical commons to train corporate algorithms."

Novara Media, July 29, 2026

The Cultural Preservation Backlash

While the story has seen low initial engagement on networks like Hacker News, the implications for librarians, archivists, and preservationists are massive. By systematically purchasing and destroying obscure, historical, or out-of-print literature, tech companies are removing physical copies from circulation entirely. It represents a modern-day digital enclosure movement, where physical knowledge is privatized, digested, and then destroyed to feed proprietary models.

If the future of machine intelligence depends on the literal destruction of our physical history, we have reached a bizarre inflection point. AI companies are no longer just indexing the world's information—they are burning the library to build the model.

This article was ultrathought.

Stay ahead of AI

Get breaking news, funding rounds, and analysis delivered to your inbox. Free forever.

Related stories