More Than Just Objects: Australian Booksellers Warn of Rare Books for AI

The process of destroying physical books to feed AI models might sound like something from a dystopian novel, but it is a very real practice known as “destructive scanning.” It involves buying books, slicing off their spines to rapidly scan every page, and then pulping the remains. This method of sourcing rare books ai training data came to light during a copyright lawsuit against Anthropic, and it has starkly divided the publishing and tech worlds.

While a US judge has ruled destructive scanning lawful under the principle of “transformative use,” that legal clarity does not extend everywhere. For Australian secondhand booksellers, the issue has become urgent. They report receiving suspicious bulk orders from a Canadian company called Zoom Books, raising profound concerns that rare and culturally significant books are being treated as disposable raw material for generative AI data sourcing. The alarm is not just about losing objects, but about losing irreplaceable pieces of history to the relentless demand for AI training data.

What Is Destructive Scanning and Why Is It Used for AI Training?

To feed massive datasets, some AI companies resort to physically destroying books after scanning them—a process that raises ethical and legal questions. Here’s how that works, and why it matters for the future of rare books. AI training data collection requires enormous amounts of high-quality text, and for many rare or out-of-print titles, a digital version simply doesn’t exist. That’s where destructive scanning comes in.

Rare books ai - real-life example
Bild: holmespj / Pixabay

The Process: From Book to Pixels

Destructive scanning is a method to obtain high-quality text data from physical books. It involves buying a copy, then slicing off the spine to separate the pages. Those pages are fed through a high-speed scanner, capturing every word. After scanning, the book is discarded—typically pulped. The result is a clean, digital text corpus that can be used for machine learning datasets. AI companies argue this is necessary to access rare or out-of-print texts not available digitally. While controversial, a US judge ruled destructive scanning lawful under copyright as ‘transformative use’, meaning the new digital copy serves a different purpose than the original book.

Why Not Just Buy Digital Copies?

If you’re wondering why companies don’t simply purchase existing digital versions, the answer lies in availability. Many older or niche books were never digitized, or the digital copies are locked behind paywalls or have poor quality scans. Data scraping from the open web can only go so far. For building a comprehensive machine learning dataset on historical topics, rare books ai training depends on can only be accessed through physical copies. So destructive scanning becomes the most efficient book scanning method—even if it means sacrificing the original object. The tension between preserving history and feeding AI’s appetite for data is only growing, and it’s a debate you’ll see more of as training demands increase.

The Australian Booksellers’ Experience: Suspicious Orders from Zoom Books

While the debate over AI training data rages on, some Australian booksellers have found themselves caught in the middle, receiving unusual bulk orders that raised immediate red flags. For independent sellers, a large order from an unknown buyer can be a welcome surprise—but not when the selection makes no sense.

Inspiration for Rare books ai
Bild: Zachtleven / Pixabay

Delfina Manor’s May Shipment

In May, Delfina Manor received three full boxes of orders from a Canadian company called Zoom Books. The sheer volume was unusual, but what really stood out was the lack of price sensitivity. The buyer didn’t haggle, didn’t ask about condition, and didn’t seem to care about the value of individual titles. For anyone familiar with the secondhand book trade, that behavior is a clear warning sign.

Sainsbury’s Books: A Pattern of Randomness

John Sainsbury, of Sainsbury’s Books, confirmed a similar experience. He noticed a slew of random, price-insensitive orders coming in from Zoom Books. The titles had no thematic connection whatsoever. One order included a 1970s soil mechanics manual, a 2010 cycling book, a 1982 poetry collection, and a local history of Hawthorn. This is not the kind of mix a collector or a library would request. Sainsbury’s Books ended up shipping multiple boxes overseas to Zoom Books, with no clear explanation of what the books were actually for.

Interestingly, high-end specialist sellers did not appear to have been approached by Zoom Books. That suggests the company wasn’t after rare or valuable editions. Instead, it targeted generalist booksellers with common, widely available stock. The pattern points to a bulk acquisition strategy, one that fits neatly with the idea of feeding a rare books ai training pipeline—where quantity matters far more than quality or provenance.

For Australian booksellers, these suspicious book orders from a Canadian book recycler have become a cautionary tale. If you’re selling secondhand books online, it’s worth paying attention to who is buying and why. A price-insensitive buyer with a random wishlist may not be a collector—they could be feeding data into an AI model, and your books might never be read again.

Legal and Ethical Questions: Is Destructive Scanning Legal in Australia?

If you’re a bookseller, spotting a suspicious buyer is one thing, but the broader legal questions around destructive scanning are another matter entirely. While a US judge has ruled that digitizing copyrighted books for AI training is lawful under the “transformative use” exception, that decision only applies within the United States. Other countries, including Australia, have not yet addressed similar cases, leaving a significant legal gray area.

The US Ruling and Its Limits

The US ruling is a landmark, but it’s not a global precedent. Australian copyright law does not have a clear “transformative use” exception like the US fair use doctrine. Instead, Australia relies on a more limited set of statutory exceptions, such as for research or study. Whether destructive scanning of rare books for AI training would qualify under these exceptions is far from certain. Until a court in Australia weighs in, the legality of the practice remains unclear. This uncertainty adds risk for both vendors and AI companies operating in the country.

Ethical Arguments: Knowledge vs. Artifact

Beyond the legal questions, ethical concerns are mounting. The practice of destroying rare books to extract data for AI models raises serious issues around cultural heritage. These aren’t just objects—they are often unique artifacts with historical and scholarly value. Vendors and rare books experts are worried about the irreversible loss of such items. The commodification of knowledge for AI training, where the physical book is sacrificed for digital data, feels like a poor trade-off to many. You might ask: Does the potential benefit of training an AI model outweigh the destruction of a piece of cultural heritage? This tension between knowledge extraction and preservation lies at the heart of the debate over rare books AI ethics. The current lack of legal clarity in Australia only amplifies these concerns, leaving the fate of valuable artifacts in a precarious position.

Scale and Financial Incentives: How AI Companies Source Books

Beyond the ethical dilemmas, a practical question emerges: how do AI companies actually get their hands on so many books? The business model behind this practice relies on intermediaries, and one name that keeps surfacing is Zoom Books. Understanding its role helps explain the scale of rare books AI sourcing.

Ideas around Rare books ai
Bild: Monam / Pixabay

Zoom Books: Recycler or AI Data Broker?

Zoom Books officially describes itself as a book recycler. On the surface, that sounds like a straightforward operation—taking unwanted books and giving them a second life. But the nature of its orders tells a different story. John Sainsbury, a prominent Australian bookseller, confirmed a slew of price-insensitive, random orders from Zoom Books. When a buyer doesn’t care about cost or content, it raises a clear red flag. This behavior doesn’t align with typical recycling; it points to a high-value end use, likely feeding data into generative AI training.

How Much Do AI Companies Pay for Books?

The financial incentives here are key. If you’re running an AI data sourcing operation, you need massive volumes of text. Buying books by the pallet from recyclers is cheaper than licensing digital content or paying for new acquisitions. The book destruction business model works because the end product—trained AI models—can be extremely profitable. Orders that are price-insensitive suggest that the buyer values the source material far more than the cost of the physical book itself. This disconnect is what fuels the Zoom Books supply chain and similar intermediaries.

Related reading: our post Surfshark vs Private Internet Access: 7 Key VPN Differences offers more practical ideas on this.

As for scale, the exact numbers are hard to pin down. The practice is understood to be widespread among generative AI companies, but no one has publicly disclosed the full scope. What is clear is that the generative AI training costs involved in bulk book purchasing are negligible compared to the potential returns. This economic reality makes it likely that the practice will continue unless legal or ethical barriers are put in place.

Protecting Rare and Unique Books: What Can Booksellers Do?

Booksellers like Tim White and Nick Dawes are wary of selling unique copies that could be destroyed, but they lack clear guidelines to identify AI-related orders. If you run a bookshop, how do you handle this ethical and practical challenge without a clear industry-wide framework?

The Bookseller’s Dilemma: Selling vs. Preserving

The dilemma is a deeply personal one. Tim White, who runs the independent shop Books for Cooks, strongly believes that a book’s value lies not just in its text, but in its physical existence. For him, the decision to sell hinges on the book’s rarity. “If a book is common, I wouldn’t cry,” White says, “but if it’s a one-off, it’s horrific.” This simple distinction between common stock and unique copies forms the core of the ethical dilemma surrounding rare books AI data orders.

Nick Dawes of Grant’s Bookshop shares a similarly conflicted perspective. He admits he would feel unhappy yet financially satisfied if someone bought his entire stock to cut up, but he draws a firm line at truly unique items. Dawes states he would not knowingly send a rare, one-of-a-kind copy to be destroyed for data. While high-end specialist dealers might not be the primary targets of these bulk purchasing orders, the risk to rare and distinctive volumes within general stock is very real.

How to Spot an AI Data Order

The core problem is the lack of clear guidelines for identifying AI orders. A large purchase of common textbooks or recent bestsellers might look like a standard bulk deal, making it hard to know the buyer’s true intent. Booksellers urgently need better tools and signals to spot these orders before they result in the destruction of culturally significant works. The secondhand book trade safeguards that have evolved over decades don’t yet account for this specific threat.

So, what can you do to protect your stock? Developing a strong internal policy for unique book protection is a vital first step. This means clearly cataloging which items in your inventory are rare, signed, or otherwise irreplaceable. When a bulk order comes in, asking direct questions about the intended use of the books can provide crucial clarity. Is the buyer a library, a collector, or a data firm? Building a network with other booksellers to share information about suspicious buyers can also act as a powerful safeguard.

Ultimately, the responsibility currently rests on individual bookseller ethics. The choice to prioritize rare book preservation over a quick sale defines the character of the trade. As the demand for data grows, the bookselling community must develop its own standards to ensure that the objects they love are not lost to an invisible digital machine. Without these safeguards, the unique and the rare will remain vulnerable.

Frequently Asked Questions

How can booksellers identify if an order is for AI training and avoid selling unique copies?

Watch for red flags such as orders for many identical titles, requests for books in pristine condition, or buyers who dodge questions about how the books will be used. Ask directly about the purpose, and verify whether the customer is a known AI company or a reseller. If you suspect the order involves rare books AI training, decline the sale or offer a digital scan instead of the physical copy.

Is it legal to destroy rare books for AI data in Australia or other countries?

Destructive scanning is the practice of cutting or dismantling books to scan pages quickly for AI training data. Its legality depends on the country and on copyright status. In Australia, owning a physical book does not give you the right to reproduce its content, but destroying a book you own is generally not illegal unless heritage protections apply. Many other countries have similar legal gaps, which leaves destructive scanning in a gray area.

Are rare or unique books at risk, or is the practice limited to common books?

Rare and unique books are the primary targets because their text is often absent from existing AI training datasets. Common books that are already digitized offer little new value to AI companies, so they are rarely part of these orders. If you collect or sell rare books, watch for bulk orders and ask about the buyer’s intended use before letting irreplaceable copies go.


Add Comment