Amazon is reportedly acquiring and destroying rare physical books to train its AI models, as these texts contain unique content not available in existing online datasets. This practice highlights the growing competition for high-quality training data as LLMs saturate public-domain and web-scraped content.
Background
As large language models exhaust publicly available text data, companies are increasingly turning to proprietary and physical texts for training corpora, raising ethical and cultural preservation concerns.
- Source
- TechCrunch
- Published
- Aug 18, 2026 at 12:38 AM
- Score
- 6.0 / 10