AI Firms Quietly Buy and Destroy Rare Books to Feed Training Data

AI developers are turning to used-book shelves for fresh training material, and a now-deleted offering from book database company ISBNdb reveals the scale: up to 1 million physical books per order, sourced from used bookstores and catalogs, with confidentiality guaranteed for buyers.

AI-generated Axo News staff avatar for Lucas Berg
5 Min Read

The practice, first reported by 404 Media, highlights a growing tension between AI training data demands and the preservation of rare, out-of-print, and specialized books that may exist in only a handful of copies worldwide.

The ISBNdb Offering

Until July 28, ISBNdb — a company that operates an online book database — advertised its ability to supply AI developers with enormous quantities of physical books. The deleted page promised books “tailored to your LLM training needs,” according to 404 Media’s reporting. The offering included older, specialized, rare, and out-of-print titles gathered from used bookstores and other catalogs.

The company also promised confidentiality to its buyers, a detail that raises questions about which AI firms are participating and how many books have already been destroyed in the process. ISBNdb has not publicly commented on the offering since the page was removed.

Why AI Companies Want Physical Books

The pivot to physical books reflects a growing problem for AI developers: the internet is increasingly polluted with AI-generated text. Training large language models on synthetic content risks degrading model quality — a phenomenon researchers have begun documenting as “model collapse.” Physical books, especially older and out-of-print titles, offer a reservoir of human-written text that has never appeared online.

Rare books and specialized academic texts are particularly valuable to AI training pipelines because they contain niche knowledge — technical manuals, regional histories, scientific monographs — that may not exist in any digital form. For AI companies racing to differentiate their models, that untapped corpus represents a competitive edge.

The Preservation Problem

The destruction of rare books for AI training data sits at the center of an ethical dilemma that the industry has not publicly addressed. Unlike digital scanning — which preserves the original — the training-data pipeline typically requires books to be digitized and then discarded or pulped. For titles that exist in only a few physical copies, that process is irreversible.

Used bookstores, estate sales, and liquidation catalogs often hold the last remaining copies of out-of-print works. When those copies are bought in bulk and destroyed for AI training data, the loss is permanent. Libraries and archives have spent decades working to preserve rare texts; the emerging book-buying market for AI training threatens to undermine that effort by removing copies from circulation before preservationists can acquire them.

A Market Built on Opacity

ISBNdb’s promise of confidentiality is not unusual in the AI training data industry. The market for training corpora operates largely in the shadows, with intermediaries sourcing text from across the web — and now from physical bookshelves — without public disclosure. That opacity makes it nearly impossible to track which books are being destroyed, which AI companies are buying them, and how many rare titles have already been lost.

The lack of transparency also complicates legal questions. Copyright law around training data remains unsettled, and physical books that are legally purchased can technically be digitized under certain interpretations of fair use. But rare and out-of-print titles may involve orphan works — books whose rights holders are difficult or impossible to locate — adding another layer of legal ambiguity to the practice.

What Happens Next

Expect scrutiny of the AI training data supply chain to intensify as more details about physical book sourcing emerge. Preservation advocates and rare book dealers are likely to push for disclosure requirements or restrictions on bulk purchases of out-of-print titles. Regulators in the European Union, already moving on AI training data transparency through the AI Act, could extend oversight to physical media sourcing.

For AI companies, the calculus is straightforward: human-written text is becoming scarcer online, and physical books represent one of the largest remaining untapped reservoirs. Whether the industry can access that reservoir without destroying culturally significant works in the process is a question that ISBNdb’s deleted page has now thrust into public view. Watch for rare book markets to tighten as institutional buyers compete with AI intermediaries for the same shrinking inventory.

— Lucas Berg, culture desk, AXO News

Share This Article