AI companies are reportedly shredding millions of books after using them to train AI models — tech giants outsource to middlemen to secretly buy up books for training material
Having contributed to the growing shortage of memory and storage, AI companies seemingly have a new target in their sights: humanity's literary history. A recent investigative report from 404 Media reveals that these companies are reportedly purchasing millions of secondhand books through intermediaries to source high-quality training data for their AI models, avoiding public backlash.
AI relies on vast amounts of data to advance, but not just any data. It has to be high-quality data. The problem is that mediocre AI-generated content, commonly referred to as "AI slop," has proliferated across the Internet. This type of content contaminates the data pool and is counterproductive for AI to train on. As a result, leading AI companies have turned to human-authored sources for knowledge, specifically print sources that predate 2022 and are more likely to contain original, uncontaminated content.
There is precedent for AI companies turning to physical books for training AI. For instance, Anthropic, one of the leading AI companies involved in a lawsuit, reportedly invested millions of dollars in extracting information from countless printed books to build its Claude AI models and then destroying them. The company bought books from Better World Books. Although the court decision affirmed that using books for AI training falls under fair use in copyright law, Anthropic faced a staggering $1.5 billion fine for maintaining a repository of seven million pirated books that infringed the copyrights of authors and publishers. Similarly, a coalition of publishers recently filed a lawsuit against Google, accusing the tech giant of allegedly and illegally using millions of copyrighted books to develop its Gemini AI models.
ISBNdb, an online database that reportedly has over 111 million cataloged books, has been a long-favorite platform for booksellers, libraries, and distributors to sell books. With the explosion of the AI industry, ISBNdb has pivoted its business to offer specialized services to bulk-purchase books for AI companies. According to 404 Media, the orders range from 1,000 copies to as many as one million books in a single transaction.
One professional bookseller, who wanted to remain anonymous, purportedly spoke to 404 Media about the unprecedented surge in book sales, which began in April of this year. The seller previously moved around 20 books in a good week, but in recent months, weekly sales have skyrocketed to several hundred books. It represents a fivefold increase over the normal volume. Other booksellers on platforms such as Alibris and Biblio have reported similar spikes in bulk purchases.
While there is no concrete proof that ISBNdb or some other AI company is making the purchase, there are some red flags. Notably, the large-scale purchases only included books with an International Standard Book Number (ISBN), the unique 13-digit code used globally to identify books. There were no patterns in terms of subject, genre, or author. It also appeared that the purchasers disregarded the pricing for the books and snapped up titles at any cost, even if they were overpriced.
During the Anthropic lawsuit, Tom Harvey, who previously participated in the creation of Google Books before leading Anthropic's "Project Panama" digitalization project, confirmed that the AI firm hired several document scanning companies. Datamation Information Services, which offers high-volume, non-destructive, and destructive book scanning services, was one of them. The former method employs different tools, like overhead scanners, flatbed scanners, or V-shaped imaging systems. The latter method, on the other hand, would have personnel gut the books and feed the individual pages into a high-speed industrial scanner. Logically, AI companies opt for the destructive route since it is more efficient and lower-cost. The result is the destruction of millions of books.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
Obviously, printed books represent a treasure trove of information for AI. However, many debate the ethics of removing books from circulation since it is uncertain whether AI companies filter the rare or even out-of-print books from the common titles during digitalization. The other major issue is that scanned books go directly into a private database to train AI, which the general public does not have access to. True, we will have smarter AI, but at the cost of the information not being available to future generations.
Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.

Zhiye Liu is a news editor, memory reviewer, and SSD tester at Tom’s Hardware. Although he loves everything that’s hardware, he has a soft spot for CPUs, GPUs, and RAM.
-
warezme Do we really want Stephen King as the basis for training AI without developing context into the AI?Reply -
Jabberwocky79 Replywarezme said:Do we really want Stephen King as the basis for training AI without developing context into the AI?
:ROFLMAO: -
SkyBill40 I don't understand why they'd destroy the books. :/Reply
This situation reminds me of a t-shirt I saw recently: "Make Orwell fiction again." -
salgado18 Reply
Inventory space for unused assets. That way, they don't even need a room to store them, and it gets cheaper. Always money.SkyBill40 said:I don't understand why they'd destroy the books. :/
This situation reminds me of a t-shirt I saw recently: "Make Orwell fiction again."
Well, sellers could create a different type of business: lend the book, then do a reverse-logistics to receive it back. Charge more for it, as it seems they don't care about costs. Book sellers keep inventories while turning a profit, preserve hard to replace books, AI companies don't look so evil, and everybody wins. I guess it's too late for that, but still, would solve many problems. -
drea.drechsler Reply
Actually: yes. AI's have to know what "popular culture" is to comment on it at all.warezme said:Do we really want Stephen King as the basis for training AI without developing context into the AI?
What is REALLY funny is ask an AI about most any a Stephen King novel and it will agree it's pure trash, even if entertaining, or great literature. It all depends on how you ask the question. -
drea.drechsler Reply
I see nothing wrong with this.SkyBill40 said:I don't understand why they'd destroy the books. :/
You're allowed to buy a book and keep it for reference without any violation of copyright laws. What you're not allowed to do is make copies except for personal use. An AI does the same thing, but it has to have the 'book' digitized in a database...that's what its 'memory' is after all... in order to reference it and make fair-use quotes for servicing 'clients'.
That digitized copy is probably a copyright violation that will get them in trouble. What I don't know is whether or not destroying the original copy of copyrighted material they legally purchased relieves them of the copyright violation for making digital copies (I am not a lawyer). Especially to use as AI's do.
Also, look at the recent Anthropic settlement so see what's at play here. There were at least two parts: one is they got pirated material (never good) but the part that got the 1.5B dollars is they kept copies (digitized in the AI's memory/database) and used them.
But at any rate it's no different from me reading the latest NY Times best seller on summer vaca then ripping pages out to start a fire in the cabin stove this winter. My library will not take donated books, they'd just toss them in a dumpster; they only want books they obtain and are in line with their collection management policies. -
bigdragon Humans write books on computers. Publishers buy dead trees to print their books on. AI companies buy up books to scan into their computer models. Then AI companies destroy the books made out of dead trees... Is there a way to do this while skipping the whole dead tree part?Reply
At least they're using the books and not the movies. I would hate to see what a Stephen King-inspired meatball recipe looks like!warezme said:Do we really want Stephen King as the basis for training AI without developing context into the AI?