Posts report bulk rare-book orders tied to AI training rights debate
Fabian Stelzer relayed a publisher’s report of price-insensitive orders for obscure books. Aakash Gupta said ISBNdb brokers anonymous bulk buys for AI labs that digitize and destroy copies under NDAs.

TL;DR
- Obscure print books are being treated like scarce training fuel: fabianstelzer relayed price-insensitive orders for odd long-tail titles, and LLMJunky said pre-2022 books look especially valuable because they predate the genAI web.
- ISBNdb openly markets physical-book sourcing for AI training, with its bulk sourcing page offering up to 1 million titles per order, while aakashgupta claimed the broker uses NDAs around anonymous bulk purchases.
- The legal incentive comes from Bartz v. Anthropic: the court order said Anthropic bought millions of print books, cut them apart, scanned them, discarded the paper originals, and got fair-use protection for the purchased-book digitization, which danshipper's reply compressed into one line.
- The backlash split into preservation and ownership arguments: danshipper said rare books should not be treated this way, hckmstrrahul called it destruction of ownership, and LinusEkenstam argued AI companies should build libraries instead.
ISBNdb's own pitch says the "world's best AI training data is sitting on a shelf" and offers printed book sourcing at scale. The Bartz order describes Anthropic's pipeline in court-record detail: purchase, strip bindings, scan, discard. 404 Media framed the demand as a hunt for pre-slop human text, while the Authors Guild says Anthropic's separate pirated-book settlement received final approval this month.
The weird orders
fabianstelzer said a small publisher had seen "insane, completely price-insensitive orders" for odd long-tail books, including examples as mundane as old printer setup guides.
When someone pointed to a different, benign-sounding scanning project, fabianstelzer's reply narrowed the claim: "Cool project but it’s not that."
The reported demand pattern is brutally simple:
- obscure titles;
- long-tail topics;
- price-insensitive orders;
- likely preference for pre-2022 text, according to LLMJunky's follow-up.
The thread also turned into a meme fast. That usually happens after the useful bit has already escaped the original context.
ISBNdb's bulk-book pitch
ISBNdb's physical-books-for-AI page says older, rare, and specialist volumes "simply do not exist in any dataset" and calls physical acquisition "the only path." Its offer is procurement infrastructure, not a scanner: source the books, consolidate vendors, ship to the customer's digitization pipeline.
The page lists the mechanics plainly:
- bulk sourcing up to 1 million titles per order;
- filtering by ISBN list, category, quantity, subject, genre, year, language, edition, or author;
- pipeline-ready delivery for scanning or digitization;
- global sourcing from hundreds of millions of titles.
aakashgupta claimed ISBNdb brokers anonymous purchases for AI labs, offers strict NDAs, and coaches buyers to describe shredding as digital preservation. ISBNdb's public LLM dataset page separately advertises 111 million metadata records plus printed-book sourcing for AI labs and research teams.
The one-for-one scan logic
The core legal twist is the one-for-one copy theory. If the physical book survives after scanning, the buyer has two copies from one purchase; if the book is destroyed, the digital scan can be framed as replacing the purchased object.
The Bartz v. Anthropic order says Anthropic spent millions of dollars buying millions of print books, stripped bindings, cut pages to size, scanned them into machine-readable PDFs, and discarded the paper originals. Judge William Alsup summarized the split result: training on the books was "exceedingly transformative," and digitizing legally purchased print books was also fair use.
That ruling explains danshipper's reply: "Legally it’s the only way it stays fair use."
Rare-book optics
The strongest objection was not about OCR technique. It was about destroying scarce physical culture for private model weights.
LinusEkenstam argued AI companies should build a physical library "that rivals the Library of Alexandria" instead of shredding rare books. His attached screenshot quoted Elon Musk saying he had asked the SpaceXAI team to preserve rare books and scan them "the hard way" rather than cut off the spine.
fabianstelzer later separated two questions: his clarification said fears about books being destroyed during scanning were likely overblown as a general topic, but "token greed" still meant even weird books were being gobbled up.
On Hacker News, one commenter in the discussion said their own reprint operation cuts bindings but stores the pages indefinitely and keeps uncompressed scans. The same commenter said no AI company had asked to train on their scanned archive.
The $1.5B piracy split
hckmstrrahul put the creator-side question bluntly: if an AI company can convert a purchased book into training material, will it provide the raw book back on demand?
The purchased-book path now sits beside a much costlier piracy path. The Authors Guild says Judge Araceli Martínez-Olguín granted final approval on July 20, 2026, to Anthropic's $1.5 billion Bartz settlement over LibGen and PiLiMi downloads, with roughly $3,000 per work before costs and fees.
That settlement only covers past input-side claims through August 25, 2025, according to the Authors Guild. Output claims, future-conduct claims, and claims for works outside the list were preserved, and Anthropic must destroy original LibGen and PiLiMi files plus copies originating from them, subject to legal preservation obligations.