Truth that Matters. Stories that Impact

Truth that Matters. Stories that Impact

Technology

Amazon Reportedly Buying and Scanning Rare Books to Train AI Models

Amazon has been purchasing large quantities of rare books, cutting off their bindings, and scanning the text to train artificial intelligence models, according to an investigation by 404 Media.

What Happened

Reporters at 404 Media placed a tracking device inside a rare book, which subsequently travelled to an Amazon facility located in Las Vegas. The site, designated as VGT3, uses an emblem depicting a dinosaur holding a book in its claws. Once delivered, rare and out-of-print volumes reportedly have their spines removed so their pages can be scanned into digital form for machine learning systems.

When questioned about the operation, Amazon provided a statement to 404 Media confirming that the company “purchases books through commercial channels to improve the products and services customers use.”

Key Highlights

  • A tracked rare book was traced directly to an Amazon facility known as VGT3 in Las Vegas.
  • The scanning process reportedly involves cutting the spines off physical books to digitise their contents.
  • Amazon confirmed it buys books via commercial channels to enhance services and products.
  • Tech companies are seeking alternative text sources after exhausting accessible online material, with Anthropic having previously relied on pirated texts.
  • Physical texts published prior to 2022 provide guaranteed human-written data free from synthetic text.

Why This Matters

Developers of large language models (LLMs) require vast quantities of written material to train their systems. Having already gathered much of the readable text available across the open internet, technology companies face a shortage of high-quality data.

Rare and out-of-print physical books offer valuable training material because works published before 2022 carry no risk of having been generated by artificial intelligence. Ingesting text generated by other AI models creates a risk of “model collapse,” an issue where an LLM’s performance and output quality degrade after consuming too much machine-generated content.

What to Watch Next

As competition for text data continues, attention remains on how artificial intelligence developers obtain offline human text and manage the risks associated with model collapse as internet data becomes saturated with AI outputs.

Frequently Asked Questions

Why does Amazon scan physical books instead of using web text?

AI developers have largely consumed the readable text available across the internet. Physical books that are out of print or unavailable online provide a fresh supply of written data.

Why are pre-2022 publications important for AI training?

Books published prior to 2022 were written entirely by humans without the involvement of modern LLMs. Training models on human text avoids model collapse, a condition where AI output degrades from learning on AI-generated text.

How was the Las Vegas facility discovered?

404 Media embedded a tracking device within a rare book shipment, which ended its journey at Amazon’s VGT3 location in Las Vegas.

Source: Reporting by 404 Media and TechCrunch.