AI Companies Destroying Rare Books to Train Language Models
AI companies, notably Anthropic, are destroying millions of rare and second-hand books to train large language models. Antiquarian booksellers in the UK and Europe received suspicious bulk orders for obscure titles. Court documents reveal Anthropic’s ‘Project Panama’ involved buying books, slicing off spines with industrial cutters, scanning pages, and recycling the rest. The practice raises copyright concerns, with publishers and authors calling for stronger regulation. Legal battles against Meta, Nvidia, and Anthropic are ongoing.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page itself is projected from evidence records.
- Current automated evidence projection
Cross-source coverage
Wire timeline
European bookstores suspect AI firms behind suspicious bulk orders of obscure books
Independent bookstores in Europe, including Kennys Bookshop in Galway, Ireland, and shops in Berlin, have received massive online orders for thousands of obscure books that have not seen interest in years. Sellers suspect these acquisitions are made by AI companies seeking to expand training data for large language models (LLMs). The orders include odd combinations like a history of Connemara alongside a guide to Ireland's Special Savings Incentive Account, or outdated driving test manuals. This follows reports of AI firms using destructive scanners to digitize and shred millions of books. While some European nations have laws against book scanning, the German Publishers and Booksellers' Association believes the orders may be routed through local addresses as collection points for eventual bulk shipping to the US, where scanning may be permissible. The practice has sparked legal battles, with Meta, Nvidia, and Anthropic facing lawsuits over book piracy for AI training.
European Bookstores Suspect AI Firms Are Buying Obscure Books for Training Data
Independent bookstores in Europe, including in Galway, Ireland, and Berlin, Germany, are receiving suspicious bulk online orders for thousands of obscure books that have not seen interest in years. Sellers fear these acquisitions are made by AI companies seeking to expand their training datasets, potentially destroying the books after digitizing them with destructive scanners. The orders include titles like 'A History of Connemara' and 'Pass Your Driving Test, 2018 Edition,' which seem illogical for personal purchase. The German Publishers and Booksellers' Association notes that scanning books violates German copyright law and suspects the orders are routed through local addresses for eventual bulk shipping to the US, where AI training may be less restricted. The article also references lawsuits against Meta, Nvidia, and Anthropic over book piracy for AI training, highlighting ongoing legal battles over copyright and fair use.
AI Companies Destroying Rare Books to Train Language Models
AI companies are destroying millions of rare and second-hand books to feed their insatiable demand for training data, raising serious concerns for the publishing industry and copyright law. Antiquarian booksellers in the UK and Europe received suspicious orders for thousands of obscure titles from anonymous buyers. Court documents reveal that Anthropic, developer of the Claude AI model, launched 'Project Panama' to acquire millions of books, slice off their spines with industrial cutting machines, scan every page, and recycle the rest. Internal documents show executives wanted to keep the program secret. Similar orders have been linked to Canadian company Zoom Books, which denies involvement. Copyright violations are difficult to prove due to legal loopholes and varying international laws. Publishers and authors are calling for stronger regulation from national governments and the EU to prevent unauthorized use of copyrighted works in AI training.
Show 2 older updatesHide older updates
AI Companies Destroying Rare Books to Feed Language Models
AI companies are destroying millions of rare and second-hand books to train large language models, sparking controversy over copyright and the future of physical books. Antiquarian booksellers in the UK and Europe received suspicious bulk orders from anonymous buyers, later linked to AI firms. Court documents reveal Anthropic, developer of Claude, ran 'Project Panama'—buying millions of books, slicing off their spines with industrial cutters, scanning pages, and recycling the rest. Internal documents show executives wanted the project kept secret. Canadian company Zoom Books was also implicated but denied involvement. Copyright violations are hard to prove due to legal loopholes and differing international laws. Publishers and authors are calling for stronger regulation from national governments and the EU, as AI systems trained on copyrighted books now produce low-cost AI-generated titles flooding digital marketplaces.
AI Companies Destroying Rare Books to Train Language Models
AI companies are reportedly destroying millions of rare and second-hand books to train large language models, raising serious concerns about copyright infringement and the future of physical books. Antiquarian booksellers in the UK and Europe have received suspicious orders for thousands of obscure titles from anonymous buyers. Court documents reveal that Anthropic, developer of the Claude AI model, launched an internal program called 'Project Panama' that involved buying millions of second-hand books, slicing off their spines with industrial cutting machines, scanning every page, and recycling the rest. Internal documents show executives wanted to keep the project secret. The trend has sparked fresh concerns about copyright violations, which are difficult to prove due to legal loopholes and varying international laws. Publishers and authors are calling for stronger regulation from national governments and the EU to address the unauthorized use of copyrighted works in AI training datasets.