AI Firms Have Scoured Japanese Bookstores for Training Data
Large-volume acquisitions of non-fiction works suggest an aggressive push to digitize physical archives for LLMs.
Updated on Sept. 29, 2026 in Artificial Intelligence

Live Poll
Should companies be allowed to use copyrighted books to train artificial intelligence models?
Used bookstores across Japan have reported an influx of high-volume orders for non-fiction titles, with books subsequently exported for use in artificial intelligence training. Anthropic has confirmed it purchases physical books to train its Claude AI models, a process that typically involves the destruction of the original materials during scanning.
Why it matters
The systematic acquisition of printed matter highlights the ongoing industry scramble for high-quality, long-form human text to improve model reasoning and knowledge depth. This practice underscores a major shift in how AI developers are bypassing licensing hurdles by targeting massive, lower-cost secondary markets.
The process involves the bulk purchase and destruction of physical non-fiction books, including history, medicine, and philosophy texts, to facilitate data ingestion. A single Tokyo bookstore reported a sale of over 1,000 volumes, while aggregate exports to the United States exceeded 50 tons over the past year.
The players
Anthropic
An artificial intelligence startup focused on building large language models like Claude, which emphasize steerability and safety.
The details
To convert physical pages into machine-readable datasets, companies employ industrial scanning systems that break down books for rapid digitization. During this process, the structural integrity of the books is compromised, and the items are destroyed as they are processed for large-scale data ingestion. This method allows AI developers to bypass digital licensing restrictions by acquiring physical copies of non-fiction content that may not be available in authorized digital databases.
Timeline
Over the last year, a major Japanese distributor sold over 50 tons of books to the United States.
Earlier in 2026, Anthropic acknowledged purchasing physical books to train its Claude AI chatbot.
The Tech Race
This physical-first strategy follows a broader industry trend of moving beyond easily scraped web data to maintain model performance gains. As high-quality internet data becomes saturated, major labs are competing to secure proprietary physical archives that are not accessible via standard search engine crawls.
Readers may see a tightening supply or rising prices for specific non-fiction and academic texts in the secondary market as high-volume institutional buyers enter the space. This trend effectively transforms physical bookstores into industrial input suppliers for AI training pipelines.
The takeaway
The transformation of Japanese bookstores into industrial warehouses highlights the lengths to which AI firms will go to secure high-fidelity training data. Watch for future updates regarding the potential implementation of new digital copyright protections that could limit the ability of AI labs to scan and destroy physical media.
Further reading
Explore the ongoing shifts in data acquisition strategies within the Artificial Intelligence sector.
Source note: This article includes information reported by Japan Today.
Live Poll
Should companies be allowed to use copyrighted books to train artificial intelligence models?







