Why tech companies are buying up tons of rare old books to train their AI models
An investigation by 404 Media found that rare books have been routed to an Amazon warehouse used for AI training data. The reporting suggests tech companies are acquiring physical archival materials to expand the text datasets used to build AI models.
Why this matters: Books in archives and rare collections were not written to feed corporate AI systems. Authors, publishers, and the institutions that preserved these works had no say in this use. If companies are buying physical copies specifically to digitize and train on them, that sidesteps the copyright fights happening in court over digital texts. It is the same underlying question dressed in different clothes: who owns the right to turn human creative work into AI capability, and does anyone have to ask permission first.
Who should care: Lawyers · Privacy officers · Compliance · General readers · AI governance · Policy
This summary is AI-assisted and may contain errors. It is an original briefing to help you gauge significance quickly — not a reproduction of the source. Always read the linked original before relying on it. See our methodology.