Back to InsightsAI

The Great Data Scramble: Why AI Companies Are Buying Old Books

By James HuangAugust 1, 2026·Updated Aug 3, 20266 min read
AI Generated Cover for: The Great Data Scramble: Why AI Companies Are Buying Old Books

The Great Data Scramble: Why AI Companies Are Buying Old Books

TL;DR: AI companies are panic-buying physical books published before 2022. Not for nostalgia — for survival. The internet is now a photocopy of a photocopy, and every generation of AI trained on AI output gets dumber. Pre-2022 books are the last provably human-written text corpus at scale. Anthropic spent millions on "Project Panama" — bulk-buying, spine-cutting, scanning, then destroying physical books to build a legally defensible human-data vault. A federal judge just drew the line: buy the book, destroy the original, keep it internal, and you're in the clear. Download a pirate PDF and you're sued for $1.5 billion. The real story isn't about books. It's about provenance as the new scarcity in an age of infinite, worthless content.

James here, CEO of Mercury Technology Solutions. Hong Kong — August 2026


The ISBNdb Gambit

A company called ISBNdb — a database of book metadata, exactly what it sounds like — recently started offering a new service to AI companies: bulk physical book procurement.

Need 1,000 books? They'll find them. Need 1 million? They'll clear out libraries, used bookstores, and out-of-print catalogs. Their pitch? "The world's best AI training data is sitting on shelves."

Here's the kicker: they specifically want books published before 2022.

Why that cutoff? Because November 2022 is when ChatGPT dropped. After that, the internet stopped being primarily human-written. It became a recursive echo chamber — AI writing trained on AI writing trained on AI writing. The tech industry even coined a term for it: "AI slop." (The Chinese translation is more visceral: AI泔水 — AI gutter oil.)

The Photocopy Problem

Think of it like this: you make a photocopy. Then you photocopy the photocopy. Then you photocopy that photocopy. What disappears first isn't the big bold headlines — it's the subtle details, the edge textures, the rare expressions, the minority perspectives. After enough generations, everything converges to gray mush.

AI researchers call this "model collapse." It's not theoretical — it's happening now. Every new model trained primarily on previous AI output loses fidelity. The long tail of human knowledge gets shaved off. The weird, the niche, the genuinely original — it all gets averaged out.

The synthetic data itself isn't poison. Used carefully, with human curation and real-data anchors, it can work. But here's the problem: the internet can no longer tell you what's human and what's machine. Is this article written by a person? An AI? An AI rewritten by a person? A person rewritten by an AI? The provenance chain is broken.

Which makes a physical book printed in 2021 — or 1952 — something very specific: a provably human artifact.

Anthropic, the company behind Claude, didn't start with book-buying. They started with piracy.

Early on, they downloaded millions of books from pirate sites. When copyright risk escalated, they pivoted hard. Starting in 2024, they spent millions on "Project Panama" — a program to buy physical books in bulk, cut off their spines, feed pages through high-speed scanners, digitize the contents, then destroy the physical books. The digital files stay in Anthropic's internal database, never distributed.

The name? No official explanation. The mission? Sounds like a movie tagline: "Destructively scan every book in the world."

As a book person, my gut reaction was disgust. Then I read the court ruling.

The Federal Judge's Three-Part Test

In Bartz v. Anthropic, a US federal judge broke Anthropic's actions into distinct legal categories:

Part 1 — Training on book text: Fair use, under these specific facts.

Part 2 — Buy physical book → create internal digital copy → destroy original → no distribution: Also fair use.

Part 3 — Download millions of books from pirate sites → store in central database: NOT fair use.

Same book. Same training purpose. Different legal outcome. The line isn't "did you pay for it?" — it's the full stack: legal purchase, single conversion, original destruction, internal use, no distribution.

Anthropic eventually settled the piracy claims for $1.5 billion. Not a fine — a settlement. Meaning they calculated that fighting was worse than paying. And $1.5 billion was still cheaper than negotiating individual licenses with millions of authors and publishers.

This isn't universal law. It's one federal district judge, one case, one interpretation. But it's enough to sketch a playbook. And ISBNdb saw the business model.

Why Claude Felt More Human

If you've used Claude, you might have noticed something: its writing often felt more natural, more grounded, more human than competitors. Now you know why. While others trained on the slop-filled internet, Claude was partially trained on millions of actual books — real human writing, edited, curated, published, with all the messiness and specificity that entails.

The $1.5 billion settlement bought Anthropic something beyond legal peace: it validated the market for provably human data.

The Real Scarcity Isn't Content — It's Provenance

Here's the reframe that matters:

We've spent decades thinking digital solved scarcity. Infinite copies. Zero marginal cost. Information wants to be free. And then AI made content generation virtually free. We can produce infinite text now. Infinite text that says nothing.

But a 1952 physics textbook — handwritten fonts, hand-drawn diagrams, mimeographed pages smelling of ink — carries something no AI can generate: a chain of human judgment. Someone decided this topic was worth years of work. Someone gathered materials. An editor chose what stayed. Readers voted with purchases and citations. The book is a fossil record of human attention and human decision-making.

I was at NYU's library last week. Picked up a 1952 freshman physics text. Every diagram, every font choice, every page layout — clearly hand-composed. Like a teacher who wrote a test by hand, then ran it through a mimeograph drum. Less efficient than AI-generated questions, sure. But behind it, you can see a person's thinking.

Most university library collections have zero digital footprint. The internet doesn't know they exist. Which is precisely why they're valuable.

Your Data Goldmine

This isn't just about books. Every business sitting on pre-2022 operational data owns a provenance asset.

• A fashion e-commerce company's buyer photos — real humans, real lighting, real imperfections

• A factory's production logs — real failures, real adjustments, real decisions under uncertainty

• Customer service transcripts from 2019 — real frustration, real problem-solving, real human interaction

The most expensive data won't be the most secret. Or the biggest. It'll be the data that can prove where it came from, who judged it worthy, and why it's trustworthy.

The New Value Equation

The old equation: More data = Better models.

The new equation: Provenance × Human judgment = Scarcity = Value.

AI companies aren't buying old books because paper beats pixels. They're buying proof of humanity. In a world where machines can generate infinite plausible text, the only thing that matters is whether a human actually wrote it, thought about it, cared enough to finish it.

Stop hoarding content. Start hoarding provenance.

The books on your shelf just became a strategic asset. The data in your archives from before 2022? That's not legacy — that's humanity's last verifiable fingerprint.


Mercury Technology Solutions: Accelerate Digitality.

Originally published on MTS Blog & Research