All news
airegulationlegal

Amazon Scans & Destroys Bulk Books for AI Training

31 Aug 2026

What happened

An investigation has found that Amazon is purchasing large quantities of physical books, scanning them, and destroying them in the process—apparently to generate training data for its AI models. The practice came to light after a bookseller, noticing a historical spike in sales over the past year, grew suspicious that AI companies were behind the surge in demand.

To test the theory, the bookseller placed an Apple AirTag inside a book before shipping it in July. The tracker led to Amazon's LAS8 warehouse in Las Vegas, Nevada, where the book was processed by an internal team known as VGT3. According to an Amazon employee, VGT3's sole job is scanning books—and as part of that process, workers cut the books' bindings, physically destroying them.

One data point from the report underscores the scale involved: a bookseller received an order of around 1,000 books through the marketplace Biblio. Separately, the report notes a Trifinity distribution center located about 30 miles south of Kenosha Regional Airport, though its connection to the tracked shipment or Amazon's broader operation remains unclear.

An Amazon spokesperson confirmed the company buys books through commercial channels, stating the purchases help "develop and improve its products and services"—without elaborating on the scanning-and-destruction process.

Amazon isn't alone in this approach. Anthropic reportedly ran a similar effort internally, dubbed "Project Panama," in which it acquired books from commercial marketplaces, cut their spines, and scanned them for training data. The rationale in both cases appears to be the same: printed books contain text that isn't readily available online and is already organized in a convenient, structured format—making them a valuable, differentiated data source for AI training.

What's still unclear

The report leaves several key questions unanswered:

  • The total volume of books Amazon has purchased and scanned to date.
  • Whether scanned content feeds Amazon's own AI models, third-party models, or both.
  • Any compensation beyond standard sale prices for booksellers or copyright holders.
  • The legal and copyright implications of buying and destroying commercially sold books for AI training.
  • How Amazon's scale compares to Anthropic's Project Panama.

The risks in play

The practice raises several concerns flagged in the investigation:

  • Permanent loss of physical books, potentially including rare or historically significant copies.
  • Copyright and IP questions around using purchased books as AI training material.
  • Reputational risk for Amazon if the sourcing is seen as a covert data-acquisition strategy rather than transparent commercial activity.
  • Market distortion in used-book pricing and availability if bulk buying continues at scale.

Why founders should care

For founders—particularly those in publishing, bookselling, data licensing, or AI infrastructure—this signals a few probable shifts worth watching:

  • It's likely that large AI companies increasingly view offline, non-digitized text (like printed books) as a premium, differentiated training-data source, which could intensify competition for physical media.
  • Booksellers may plausibly see continued demand spikes as more AI companies pursue similar acquisition strategies, creating a short-term revenue opportunity but also market unpredictability.
  • Companies with book, archive, or print-based IP should consider monitoring whether bulk-acquisition practices like this could affect their own supply chains, pricing power, or copyright exposure.
  • The emergence of dedicated scanning operations—Amazon's VGT3 and Anthropic's Project Panama—suggests physical-media processing pipelines may become a more common, if largely invisible, part of AI data infrastructure, which could create adjacent opportunities for vendors who scan, digitize, or broker book data at scale.

As with much of the AI data-sourcing landscape, transparency remains limited. Founders operating anywhere near physical content, licensing, or bookselling markets should treat this as an early signal rather than a settled trend, and watch for regulatory or legal clarity on how commercially purchased books can be used for AI training.

Sources