AI companies are purchasing rare hardcover books and destroying them to extract text for training datasets, raising concerns among book preservation advocates. The practice highlights growing tensions between AI data acquisition and cultural heritage conservation.
Background
Large language model training requires massive amounts of text data, leading some AI companies to seek out physical books as data sources through scanning and OCR processes.
- Source
- Good e-Reader
- Published
- Aug 20, 2026 at 04:44 AM
- Score
- 5.0 / 10