Secret tracking device placed in rare book ends up in Amazon processing facility — destroying books to train AI models is 'all' the Vegas warehouse does

A book with its spine strained.
(Image credit: Getty Images)

AI companies are reportedly buying up millions of books, some of which are rare, and using them to train AI models, often destroying the book in the process. And we now have more specifics about how this process works, thanks to an investigation from 404 Media. The outlet placed a tracking device inside a rare book that was part of a large, 1,000-book order — and it ended up in a Las Vegas Amazon warehouse.

Employees who spoke to 404 Media said they spend their time at the facility (named VGT3) cutting the spines off of books and feeding them into a scanner. VGT3 is reportedly part of another, larger facility in Las Vegas called LAS8. Its purpose, according to the employees who spoke to 404 media, is solely to cut the spines off books and scan them. "All we do is scan books," one employee told the outlet.

A logo of the Amazon VGT3 facility.

(Image credit: 404 Media)

If there was any doubt about the purpose of the facility, a logo with VGT3 shows a dinosaur holding a book, looking like it's about to rip it up and take a bite. At the very least, the dinosaur isn't reading the book. You can see that image captured on 404's website above.

Latest Videos FromTom's Hardware

The investigation started in July, when 404 asked a bookseller if the outlet could place an Apple AirTag in a book inside a 1,000-book shipment. A shipment this large is unusual, though they've reportedly become increasingly common in the AI era. Book sales are up, and many suspect that isn't due to a rise in readership, but for training frontier AI models. The order was submitted through Biblio, an online marketplace that claims to be the "largest independent book marketplace in the world."

Often, these large orders contain a scattershot of different books. An independent bookseller in Ireland, for example, received an order for 5,000 books that they suspected was for training AI. “Some was high-quality non-fiction, like A History of Connemara, and the next thing might be The Eddie Hobbs Guide to your SSIA,” Tomás Kenny of Kennys Bookshop said at the time. 404 Media says that employees at VGT3 are required to scan the ISBN (barcode) of each book, lending credibility to the theory that AI companies are working through a list of every book that has ever been published.

Or, at least, every book with an ISBN. A bookseller told 404 Media that these large orders "never" include rare books that don't have an ISBN.

Earlier this year, court filings revealed Anthropic's 'Project Panama,' which kicked off in 2024. Documents as part of those filings make the project's purpose clear: "Project Panama is our effort to destructively scan all the books in the world," said an internal planning document. A judge ruled that Anthropic was allowed to use books to train its AI models. However, it was fined $1.5 billion for keeping 7 million pirated books in a central library. In a 2025 lawsuit against Meta, it was revealed that it had pirated nearly 82TB of books.

Scanning books en masse is nothing new. In 2005, Google spearheaded Google Books by scanning out-of-copyright titles from an on-campus library using specialized scanners. The books were then returned. Presumably, Google made use of some type of V-shaped scanner, which lays a book out naturally so as not to disturb the spine and binding.

In the case of VGT3, 404 reports that employees cut off the spine of the book before scanning, presumably to feed the pages flat into an industrial scanner. The insatiable hunger for data in frontier AI models seems to be moving at a faster pace than Google's early book digitization efforts.

It's hard to say why a facility like VGT3 operates in this way, though it likely comes down to cost. As major companies like Meta and Anthropic have been caught with a library of pirated books, they now need to buy them. And when purchasing thousands of books at a time, it's probably much cheaper to get secondhand copies from marketplaces like Biblio than it is to spend full price on digital versions of those books (if digital versions exist in the first place).

Google Preferred Source

Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.

TOPICS
Jake Roach
Senior Analyst, CPUs

Jake Roach is the Senior CPU Analyst at Tom’s Hardware, writing reviews, news, and features about the latest consumer and workstation processors.

  • jabliese
    "Books by scanning out-of-copyright titles from an on-campus library using specialized scanners."

    Better would be "Books by scanning public domain titles ..."
    Reply
  • GoofyOne
    I have a question ... what if the books and documents they scan into their database are full of false, incorrect, and outdated information???

    Eventually the so called AI just becomes giant database of useless drivel.

    The AI terminology is being overused and misused, it's mainly just a giant database with a fancy parser at the front end, and a inference engine that returns what it thinks you want to know, or rather tells people what it wants them to think.

    I an sure they can be useful for some things, stupid people obviously think they are great, but it's just fake pseudo intelligence coded by humans. Remember who it was that created and coded these things ... humans!!! Therefore humans are much smarter.

    GoofyOne's 2c worth which may, or may not be, actually worth 2c.
    Reply
  • PEnns
    Another reminder how benign AI is, how it is really used for our betterment and is in the best interest of humanity..... to have rare books destroyed in its training.

    One day we will have libraries stuffed with nothing but cartoons and carppy books but hey, we have AI that can create all kinds of slop to entertain simple minds.

    Unbelievable!
    Reply
  • USAFRet
    GoofyOne said:
    I have a question ... what if the books and documents they scan into their database are full of false, incorrect, and outdated information???
    Currently, the main source of data for the AI bots is reddit.
    Which, as we all know, is full of false, incorrect, and outdated information...;)
    Reply
  • SirJefferE
    Just thought I'd note that the second paragraph is incorrect:

    Employees who spoke to 404 Media said they spend their time at the facility (named VGT3) cutting the spines off of books and feeding them into a scanner. VGT3 is reportedly part of another, larger facility in Las Vegas called LAS8. Its purpose, according to the employees who spoke to 404 media, is solely to cut the spines off books and scan them. "All we do is scan books," one employee told the outlet.

    Those quotes were not from Amazon employees speaking to 404 media. They were from employees speaking to each other on an online forum. From the original 404 media report, relevant parts bolded for emphasis:

    but Amazon employees who work at this location and who discuss working conditions there with other Amazon employees online explain that the the north end of the LAS8 warehouse, where I saw the book arrived, housed a different Amazon operation with a different code: VGT3.

    “I work at VGT3 here in Vegas, and all we do is scan books,” one Amazon employee wrote on a forum for Amazon workers. “Some are assigned to cut books, and others go to receive where they get books and scan the bar codes. We didn't have rates, but now we do, but it's not stressful.”

    “Working at VGT3 is nice all we do is scan books,” another Amazon employee wrote. “It's so cool.”

    I saw VGT3 employees online talk about how working in this operation is a good but boring job because workers do the same, easy, repetitive tasks all day. I also saw some discussion indicating that VGT3 jobs are desirable for this reason, and because the warehouse sometimes offered night shifts.
    Reply
  • havfy
    USAFRet said:
    Currently, the main source of data for the AI bots is reddit.
    Which, as we all know, is full of false, incorrect, and outdated information...;)
    Currently, a main source for people on the Internet who don't use AI is reading comments like yours that have no basis in reality, no sources, and plenty of evidence to the contrary.

    AI makes mistakes. AI sometimes gives people wrong information. But it's much more reliable than what you get from random comments on the Internet. So fact check everything, including this.
    Reply
  • COLGeek
    havfy said:
    Currently, a main source for people on the Internet who don't use AI is reading comments like yours that have no basis in reality, no sources, and plenty of evidence to the contrary.

    AI makes mistakes. AI sometimes gives people wrong information. But it's much more reliable than what you get from random comments on the Internet. So fact check everything, including this.
    You may want to reconsider your position.

    Reply
  • havfy
    COLGeek said:
    You may want to reconsider your position.

    No, I'll still take facts over random posts on the Internet that say "LLMs like..." from which people decide that all of AI works a certain way.

    Plus, this graphic contradicts your statement. It shows that two specific LLMs have a plurality of citations for that source, but that it's not the main source. They give far more for other sources. And it doesn't say that it got most of its information from those, but that that's the rate at which they gave citations to sources. It says nothing about where they get information.

    That's why you have to be careful about what you read in random posts, because this is an example of somebody wildly misinterpreting things, extrapolating without rationale, confusing common sources that two LLMs point people to with where it gets its data (it can't say that it read 1000 books on the subject for which there's no online source, but here's a reference, for example), deciding that because two of them do, they all must do the same, and not knowing the difference between a plurality and majority. Most means more than 50%.
    Reply
  • USAFRet
    havfy said:
    No, I'll still take facts over random posts on the Internet that say "LLMs like..." from which people decide that all of AI works a certain way.

    Plus, this graphic contradicts your statement. It shows that two specific LLMs have a plurality of citations for that source, but that it's not the main source. They give far more for other sources. And it doesn't say that it got most of its information from those, but that that's the rate at which they gave citations to sources. It says nothing about where they get information.

    That's why you have to be careful about what you read in random posts, because this is an example of somebody wildly misinterpreting things, extrapolating without rationale, confusing common sources that two LLMs point people to with where it gets its data (it can't say that it read 1000 books on the subject for which there's no online source, but here's a reference, for example), deciding that because two of them do, they all must do the same, and not knowing the difference between a plurality and majority. Most means more than 50%.
    OK, maybe I should have stated it differently...
    "The largest source of facts..."

    In any case, the current consumer grade LLMs get their data from where? The general internet. And return their statements based on that.

    Unless, of course, you have some other info on where the LLMs get their data from.
    Reply
  • COLGeek
    havfy said:
    No, I'll still take facts over random posts on the Internet that say "LLMs like..." from which people decide that all of AI works a certain way.

    Plus, this graphic contradicts your statement. It shows that two specific LLMs have a plurality of citations for that source, but that it's not the main source. They give far more for other sources. And it doesn't say that it got most of its information from those, but that that's the rate at which they gave citations to sources. It says nothing about where they get information.

    That's why you have to be careful about what you read in random posts, because this is an example of somebody wildly misinterpreting things, extrapolating without rationale, confusing common sources that two LLMs point people to with where it gets its data (it can't say that it read 1000 books on the subject for which there's no online source, but here's a reference, for example), deciding that because two of them do, they all must do the same, and not knowing the difference between a plurality and majority. Most means more than 50%.
    Most used source seems pretty clear, others cite the same.

    I (and the person you quoted earlier) are far more adept than the average reader. Semantics aside, there is zero doubt (factually) that reddit is the most used source as depicted. Of course that doesn't mean exclusively reddit.

    https://peec.ai/blog/top-domains-cited-by-ai-search-analysis-based-on-30m-sources
    https://friendlychro.com/2025/08/19/where-ai-gets-its-information-2025/
    Regardless, you do you. Peace.
    Reply