AI Book Scanning: Just What Is A Rare Book?

One of the stories of the last few weeks has been that AI companies have been scanning books in very large numbers in order to train their models with content guaranteed to have been written before 2002, and thus AI free. It’s caused some outrage, because of the size of the operation, and because the scanning process is destructive. In particular the phrase being bandied around is that these are rare books, and it’s this phraseology I find problematic. I think it’s time to unpack why that is the case.

It’s Not Book Burning, Folks

Before I worked for Hackaday I had a long career in and around the publishing industry, mostly on the electronic side, but from time to time crossing paths with my colleagues in the world of paper-based publishing. I understand the appeal of a good book, I’ve spend a lot of my life among bibliophiles, and let’s just say I own a few books myself. In particular I understand the symbolism of destroying books, bringing to mind as it does the actions of repressive regimes. I have stood in Bebelplatz in Berlin where the photo of Nazi student organisation members burning the library of Magnus Hirschfeld’s institute was taken in 1933, and if you know me, you’ll have an idea why that’s close to home. But for all that, what the AI companies are doing is not the same thing.

In this case they’re destroying the books for two reasons. Firstly, as I remember from a previous employer in the publishing world, it’s much easier to digitise a stack of papers than it is a bound book. Thus I’m pretty sure that’s one reason they remove the binding before digitising the pages. Then secondly, as I understand it it’s a copyright issue. If they buy a book, digitise it, and destroy the physical copy, they can legitimately claim that only one copy of it exists, and they hope, sidestep copyright claims from publishers.

When Rare Maybe Isn’t Really Rare At All

The ISBN panel with barcode from Bil Herd's Back Into The Storm.
You’ll only see one of these numbers on a book published since 1970.

Perhaps the most pertinent question then is just what are the books being scanned and destroyed? They’re almost universally described as “rare”, but is that accurate or sensationalist? It brings to mind a dusty library filled with priceless tomes hand-transcribed by monks which it would be a crime to destroy, but there’s something that explodes that vision in an instant.

If you read the reports of what’s happening, they are ordering books by ISBN number. That’s an international system for identifying books, which was only introduced in 1970. If they’re ordering a book by its ISBN, it’s no medieval illuminated manuscript.

So the books being scanned and destroyed are relatively new, but can they still be described as rare? In that case just like anything else mass-produced since 1970, how many survive depends on the size of the original print run and how valued they have been since. So a few of these books can be physically rare in the sense of being uncommon, but if they are next-to-valueless, it’s fairly obvious nobody has particularly cared about their survival up to now. It’s likely that any books printed since 1970 which are both rare and of value will have their future assured, so the AI industry is not committing the wanton destruction of culture the reports would like to suggest.

You Need To Know Just How Many Books Get Pulped Every Year

A dumpster full of books
People are often shocked when they find a dumpster full of books for recycling. Ricky Shore, CC BY-NC-ND 2.0.

I’m left feeling that a combination of the symbolism of destroying books along with a distaste for AI companies has inflated the status of these books well beyond their worth. If people truly had a care for old books they would be shocked to know how many are pulped each year by the paper recycling industry.

The publishing industry has been churning out mass-produced books for centuries now, so the world is awash with old books. Where do these newly-minted bibliophiles imagine they all go, to the Great Library In The Sky? I haven’t even touched yet upon the publishing industry, which pulps vast quantities of brand new unsold books every year. Where is the outrage, I ask?

We all love to dunk on AI companies, and Heaven knows, there are plenty of reasons to do so. Among all those reasons, sadly destructively digitising books is pretty low on the list. Please, ask the other questions, the ones they really don’t want to answer!

Header image: Yair Haklai, CC BY-SA 4.0.

113 thoughts on “AI Book Scanning: Just What Is A Rare Book?

    1. Criticism as such is valid, thought. The constructive one, at least.
      Nothing is worse than indifference, people used to say.
      Critical thinking can help seeing pros and cons of things. 🙂

      1. Or understanding the issue correctly.

        The complaint isn’t about destroying some surplus books that nobody cares about. The issue is having a system that incentivizes book burning through competition for petty profits, even competing to destroy books to the last copy so their rivals would not have them.

        These books may not be rare now, but they will become rare once enough have been destroyed.

        What’s worse, it’s not burning unsold new books that nobody cares about – it’s burning books that someone bought, read, and thought important enough to re-sell to a used book store, who paid money for it because they thought it had some value. It’s burning books that have gone through the human filter saying “this should be preserved”.

        1. For once I agree with you fully, though I’d be willing to bet many of these books are rare now, and may always have been pretty rare even if they are relatively modern. Your Tom Clancy, Cussler types who hit best seller lists with most of their many many books will have had so many reprints and big runs in the first place even of their poorer novel that the book isn’t at all rare, so destroying one of those isn’t a big deal.

          But smaller more niche authors that might well have only had perhaps a few thousand books in their one and only print run, doesn’t mean people haven’t cherished and cycled those books around over the years putting real value on them, simply that they never made it big enough, likely for having a relatively niche audience to have huge volumes of their book made. For instance I have 3 history books that are really well researched, clear, full of period images for a small town of very little historical note in its own right here in the UK – just how many people are going to be interested in that?! I’d bet there were only a few hundred printed in the first place! Still a book with real value to those that have a reason to be interested though.

          1. If I were an AI company, I would stay well away from any best-seller list because those books, music, movies, etc. are already as close to AI slop as humanly possible. It’s a closed feedback loop, because people buy what you sell them – because they wouldn’t know of anything else. The bigger the publisher and the bigger the market they’re serving, the more formulaic and self-similar the books they sell become. There are whole teams of marketing experts and ghost writers and editors to plan the next best-seller by committee, and since the publishing and marketing is happening on a global level, they’re essentially forced to repeat their own output.

            It’s like observing that the single best selling food in the world is a Bic Mac with Fries, or Kraft Mac & Cheese, or Pizza with Bud Light, so that’s all you’re gonna put on the menu of your new global fast food monopoly – so creating a self-fulfilling prophecy. You can’t tell what your customers actually prefer, and neither can they because they haven’t had it – only what they will tolerate when nothing else is on offer.

            I’m reminded of asking a computer store clerk what hard drive is the best they have, and they replied “We sell a lot of Maxtor” – yeah, because at that time they put those things in cheap PCs and they broke down a lot, so customers came in to buy replacements and they were offered Maxtor again.

            Perhaps that’s why they’re going for the second hand book stores, because that’s a filter that rejects the forced fashions – the “best sellers” command little value so people don’t bother to collect and re-sell them.

          2. If you ask ChatGPT about what it thinks about scanning old books for training, it answers:

            The irony is that secondhand bookstores may function as a corrective to commercial filtering only while the books remain available for human discovery. Turning them into an unregulated raw-material supply for private AI systems risks destroying precisely the long tail that makes them culturally valuable.

            In other words, if the books become lost to human readers, the AI companies lose sight of what books people actually like to read, and the system collapses back to following just what the publishers want to sell.

        2. What’s worse, it’s not burning unsold new books that nobody cares about

          This is precisely what it is !

          – it’s burning books that someone bought, read, and thought important enough to re-sell to a used book store, who paid money for it because they thought it had some value.

          No it’s not !

          It’s burning books that have gone through the human filter saying “this should be preserved”.

          Again. No it is not.
          Books are bought by consecutive ISBN numbers, with no special consideration about if they are rare or not and with no thought to deplete them for others

          1. This is precisely what it is !

            By your own account they’re not differentiating whether the book is surplus inventory or something else. What they do get is books that mainly exist in second hand shops, because where else would the majority of the old books be? Especially if they’re looking to buy them for cheap, not buying new books just out of print.

            Picking up consecutive ISBNs is just a convenient way to order them. The orders are directed towards particular shops that might have them.

          2. The point is, the books that still exist are those that people cared about. The books that nobody bought were mostly tipped into the chipper already.

            Some of the books they buy will be new, unsold and surplus, provided that they choose ISBNs in the range of new books. The majority of the old ISBNs that can still be found on the market are those that went through the human filter that decided to keep them, and now they’re destroying them.

            When multiple companies are grabbing the books, it’s a question of who gets them first before the inventory runs out, so everyone tries to buy and scan everything first, racing to destroy books.

        3. “that someone bought, read, and thought important enough to re-sell to a used book store, who paid money for it because they thought it had some value”

          I sell books to say, HPB because I dont like them. If I liked the book and thought it important Im NOT getting rid of it. I might give or lend (rarely!) it to someone who can use or appreciate it but if its valuable to me I keep it.

          I like that books are being digitized – from a hacker POV thats great! I own many books that are very rare and valuable to me. A few I had kept my eye out for 40+ years. Im not going to relinquish them for AI as AI isnt very ‘intelligent’ at all yet. But thsts another story.

    2. That’s assuming, or at least suggesting that people only have a single reason to hate AI. Sure, they will quickly latch on to one MORE reason, but one wrong reason among dozens isn’t such a sin. Especially as the average person has little access to verifiable information and is exposed to lots of disinformation. I’m not going to harp on where or who it’s coming from, and what motivations may be involved, but, you are clearly blaming the victims here. We can reason about how AI data centers are effecting electricity bills, their water use in drought areas, and how the public will have to pay the bailouts through 401ks, etc., when the bubble bursts

    3. Yeah. It’s always annoying to try to explain to somebody that it’s unwise to use dishonest or mistaken venues of critique when there are plenty valid and proven routes to take, because you’ll just end up making the thing you’re critiquing look better by lending them the image of a moral panic or witch hunt. Note that I say “image of,” not that that’s what is truly happening. You’re doing their PR defense for them.

      It’s especially frustrating when you try to point this out and just get accused of being a shill for that thing you’re trying to more effectively critique. There needs to be a name for this phenomenon, it’s too wordy to explain each time. Anyone know what I’m talking about?

      1. Yeah. It’s always annoying to try to explain to somebody that it’s unwise to use dishonest or mistaken venues of critique when there are plenty valid and proven routes to take, because you’ll just end up making the thing you’re critiquing look better by lending them the image of a moral panic or witch hunt.

        Hasn’t the last decade of politics proved this wrong, though? The winning strategy seems to be throwing mud at a wall and seeing what sticks. Factual accuracy doesn’t win arguments or influence future events. If you hate this, take it up with god. It’s the nature of the world.

    4. The practice is about censorship and gatekeeping. If all of the books are online and under the control of AI then those that control the AI will be able to decide who gets what information and how much of it. It is a sick use of technology and should be banned. Why do you support this tyranny? Why do you hate freedom?

    1. I’d like to see all this effort put into questioning the copyright lobby.

      This isn’t a problem specific to “extremely well-funded AI companies.” These are the same laws that cause problems for more noble projects like archive.org.

      1. Dude they’ll straight up delete your comments if you start questioning copyright law. Ask me how I know. Or as they’ll explain it, they let mass-reports replace moderation because that’s just how the website works, tee-hee! Nobody knows where these mass reports come from, of course. Huge mystery.

    2. Bartz v. Anthropic (June 2025) is most likely where this came from so the courts have already ruled on the “questionable” parts. The emotional parts, not so much.

    3. Riiight, I’m sure these companies are spending thousands per rare book, just for training, when they could be spending a dollar per book for common ones.

      IMO the single biggest reason the books are for sure not rare, is companies hate spending if they don’t have to. Between a rare book not in the training data, and a common one not in the training data, they will take the cheapest every time.

      1. The real question isn’t about the rarity of these books, but what will happen to books in general after the AI companies have bought all the loose stock and then destroyed them.

        Scanning to digital and destroying the physical copies is a death sentence, because the digital data is much more volatile. A book can survive hundreds of years, flash memory erases itself in couple decades, and if the company folds they’re likely to just destroy everything anyway.

        The survival of some random book that isn’t a classic or highly regarded otherwise depends on there being millions of copies sitting around in used book stores and on peoples’ shelves. Historians in the future will not find these books if we keep burning them until they become actually rare, because there’s little point here and now to conserve them for the future.

      2. Not all rare books are insanely expensive – that niche book on model locomotive building for the hobby machinist for instance is likely a very very good book, but with only a few hundred people interested in the topic and actively seeking a new book for their hobby at a time, in many cases then passing the book on when they are finished with the project (or when their heirs sell off the workshop collection not having the bug themselves) that sort of book tends to not be all that expensive despite their relative rarity!

    4. No, not a hack, but if a Hackaday writer has relevant experience on a topic in the news, I’d like to hear it.

      Hackaday has come a long way from its original 1 hack every day. I like the long form articles and, so far, do not see any evidence that Hackaday is going down the commercialization / ens#!tification path.

      1. There is always someone crying HaD covering something that isn’t a hack. Even on subjects that are in band for a hack.
        Best to let the babies cry themselves out sometimes.

      2. Thank you. And yes, it’s my publishing industry background informing this piece, not my engineering background.

    5. because all they’ve been able to do lately is repost content thats put up on youtube and/or written by and about AI.

  1. It’s Not Book Burning, Folks

    Hm. I can’t help it but that reminds me of two famous (and overused) quotes..
    “No one has the intention of building a wall …” (GDR)
    “The pensions are safe.” (FRG)

    🥲

    PS: Just because books are casually being dumped through an established practice, does it make the situation any better?
    And if there’s a wildfire, throwing away burning cigarettes nolonger matters?

    Perhaps not everything can or should be preserved,
    but the whole idea of careless dumping/recycling maybe should be questioned these days. At least for literature (written thoughts) or art.
    Especially because physical copies are nolonger that common.
    It might also be an ethical or philosophical question to society, as well.

    LLMs or not, many “lesser” books from, say, the 80s (nerdy stuff or soap operas, comics, magazines, life style magazines etc) might be really hard to obtain by now, even in the antiquarian bookshops.
    Let’s say some Star Trek novel or something along these lines.
    Sure, they might be somehow archived in a national archive, but access to that is restricted to ordinary people.

    1. Not to mention the fact that there are many limited run books that have ISBN’s but there might only be one or two copies left in existence. For example: genealogical books containing priceless family histories that cannot be accessed any other way.

      1. Even in such a contrived case, I am happier knowing those books are being digitized. The alternative is that they inevitably rot and end up in the trash.

        Sure, I would rather it be done in the public domain rather than a private archive. But copyright law and capitalism doesn’t necessarily allow for that.

        At least in this case their digital copy will likely survive past their copyright expiration and perhaps make it into the public domain.

        1. We can hope for that, but given the track record of a lot of these companies, I would be surprised. They want to keep or sell any exclusive model training abilities. It’s in their best interests to hold data that other companies and entities do not.

        2. The digital copy is likely not going to be accessible to anybody. So it is, for all intents and purposes, lost. Digital media is also really hard to archive. Really hard. Anybody still has a 5.25″ floppy drive? 8″? IOMega zip drive? Even a CD-reader for your computer? How many of the CD-R are still readable? Did you really rotate? Can you still read the format of the data written even if you can still read the bits?
          (I guess many of us do have at least a CD or DVD reader – here… I also guess many of us have still some CD-R or CD-RW that are probably unreadable by now)

          1. Exactly! They aren’t archiving them so that they can be preserved and accessable for humanity, they’re copying them so they can rip them off and pass off derivative work as original.

            As someone who buys used books fairly regularly, this is just yet another way that rampant AI is dramatically raising costs for the peons.

          2. Very true. The digital archive can be preserved and maintained if the company cares to put in the effort, but they don’t, its probably even in their best interests for the archives to be lost so their trained “AI” becomes the sole arbiter of truth. So they probably won’t even bother to try at all!

        3. I am happier knowing those books are being digitized.

          I’m not. A book can sit on a shelf for 200 years and still be legible. Digital copies take active maintenance of the hardware and migrating to newer and newer hardware every few years – and if there’s nobody paying then the power goes out and the hardware becomes scrapped, with the data that was on it.

          Scanning to digital and then destroying the physical copy is basically ensuring that the data will be lost in the next couple decades, or sooner when the AI bubble bursts and these businesses go under.

          1. and these training sets may well have even shorter lives than just something sitting on a random persons hard drive would. They will be eaten by some corp, sold for a while then deleted when the next set comes along.

          2. Definitely. It’s easier to read a book from 1880 than a floppy disk from 1980. Of course, now someone will come along saying “But no, [current storage format] will be around forever this time! Trust me!” Just ignore them. Paper is king.

        4. Thats not a contrived case. Thats a known case that eats away at the casual attitude the article takes towards destroying information just because they don’t like to to see the individual case as important.

      2. Oddly genealogy is a bad example. It’s probably one of the most digitised, microfiched, and preserved subjects there is.

      3. genealogical books containing priceless family histories that cannot be accessed any other way.

        So what ? If even the families potentially interested in those stories discarded them so who really cares at the end ?
        Information is lost everyday. Not such a big deal.

  2. Thank you for combatting misinformation. It’s even more important to think critically about things that support our biases, since things which act against them are automatically treated critically.

  3. ” If they buy a book, digitise it, and destroy the physical copy, they can legitimately claim that only one copy of it exists, and they hope, sidestep copyright claims from publishers.”

    Haven’t we seen this before with, ReDigi, Aereo, and Zediva?

  4. …not sure it’s even mildly appropriate, here, but the first thing that came to mind was Mark Twain’s quote on the ‘Classics’–

    “A classic is a book that everyone wants to have read, but nobody wants to read.”—SLC

  5. AI companies have been scanning books in very large numbers in order to train their models with content guaranteed to have been written before 2002, and thus AI free.

    Does this indicate that the AI companies are already having problems with so much AI slop/crap on the internet? Isn’t this effectively an admission that their AI does indeed produce slop/crap? And if their own AI can’t cope with all the AI slop/crap they have produced isn’t that effectively an admission that they have compromised the usefulness of the internet?

      1. ^this
        Imagine trying to learn new things and the only stimulus you have is things you’ve already written.

        The only way it would ever allow you to improve if a massive amount of time and energy went into you writing massive quantities of things on all sorts of subjects and then a massive amount of manpower is put into giving you back your right answers.

        Its not as efficient as reading a new book by someone else.

        1. Friends, models are regularly trained on their own outputs, and have been since well before the advent of the transformer architecture that dominates today. It’s a bit absurd to say that “everything ever written” isn’t quite enough text to train a state-of-the-art model, but it’s true nonetheless, and they all rely on synthetic corpora.

          The funniest example is the account of how OpenAI’s models became obsessed with goblins: a small quirk starting in GPT5.1 or earlier, eventually reaching broad levels by continuing to train on the output of successive generations: https://openai.com/index/where-the-goblins-came-from/

        2. You do actually train yourself on your own outputs, of course. You asses whether your last attempt worked or failed and use that for your next go. We call it learning, and we do it all the time.

          This is the real limitation of the LLM strategy — that it doesn’t learn beyond the training phase.

          Buying books is an attempt to continue to enlarge the training dataset, and since the easy/cheap stuff is already ingested, they have to move on.

          It absolutely makes me wonder what the marginal benefit is of adding a fringey book to the training set. But with so many parameters, all my intuition is out the window anyway, so you might as well try it and see. My guess is that’s what they’re doing.

          1. This is the real limitation of the LLM strategy — that it doesn’t learn beyond the training phase.

            It’s rather that the LLM doesn’t have a goal in its learning. It isn’t trying to write a good book by any preference or quality metric of its own, it’s merely happy to replicate whatever it was given, so even if it did learn from its own output it would simply degenerate.

    1. Yes, a large fraction of the text on the modern internet is AI generated. And for some use cases, it’s fine to train on that. You would essentially be creating a “distilled” model which is somewhat limited by the quality of the ancestor models.

      If you want to create a frontier model which surpasses previous models, you may wish to exclude outputs from those previous model from your training.

      I also mourn the loss of the human internet, but this practice would pop up with or without slop.

  6. The article is well written, as some people read ‘rare’ as destroying one-of-a-kind books. I also read a few issues in the text, though. That a book has an ISBN does not mean it is mass produced. Some books only have a hundred copies sold, yet it is published. A second claim is that no-one cares about these books is also not true. I’m a collector of specific tech and security books, and have a list of around a hundred book which I’ve never seen for sale at a reasonable price. ‘Genius of British Locks and Lockmakers’ for example, ISBN 9780957491328. This book is written by Tony Beck the chair of The Lock Collectors Association in the UK, and his writing is great. If you like British locks, that is. If you have one available, please reach out.

    It is unlikely the AI company will buy a copy of every book, but the more of these ‘rare’ books are scanned, the more difficult they become to find. And there are several dozen companies doing this simultaneously, and more will follow. Possibly a bigger issue, what will these companies do when they are finished eating one of every book? Where does this data hunger stop? All I expect is that it won’t stop with books. And quickly comes for every other source of knowledge.

    One good outcome would be if these scanned books were shared with the archive.org, but at this point I would consider it form of greenwashing. Even if it were a welcome one.

    1. It’s also useful to consider that they want any competitive edge they can get. It’s better for them to have digitized the only copies left of whatever information they can get their hands on. I would not be remotely surprised if some are intentionally buying all copies of a book, digitizing one, and destroying the rest, although they would certainly work hard to keep that a secret for PR reasons. It’s just unfortunately in line with the rest of their ethos.

      I appreciate this article, it’s always good to try get the proper perspective and correct misconceptions, but it’s hard not to take the most pessimistic view, given the track record of the companies and the people and politics involved. A lot of the AI evangelist culture has overlap with the past NFT craze, and many of those people were willing to destroy the original real objects in order to give their blockchain URLs more value.

          1. There is a requirement in American book publishing. It’s very specific and negates a substantial amount of the bullshit arguing. It is however limited to US jurisdiction.

            Since nobody has brought it up I just assume there aren’t any experts in the room.

      1. Nah, that’s rubbish. No piece of information or particular work is critically valuable to an AI training workflow, only volume and variety matters. So trying to prevent other entities from having access to something specific by buying all copies of it is both totally infeasible and entirely pointless.

        Especially since anything actually valuable for it’s specific content is already either held by interested parties who won’t sell, or has a digital version which can multiply infinitely. And anything that isn’t valuable isn’t worth deliberately eliminating.

        1. You wouldn’t believe how many people are out there destroying just for the sake of destroying, especially if it is something of value for someone else. This AI/book thing is just a symptom.
          How much “neutral” or “good” destruction do you have to do to get away with some really harmful destruction without people stopping you, or even noticing?

        2. Maybe. I think it would be rubbish in a sane world run by sane people. A lot of them literally do believe they’re creating their own god that will make them immortal and eventually spread its power over the entire galaxy, though.

          They are spending hundreds of billions of dollars to the point that AI training, infrastructure, etc outpaced consumer spending in GDP growth in the U.S. last year. Some of that is circular spending and not sane accounting, but still, it’s vast amounts of money.

          The book purchasing and management has to be algorithmic, and is likely assigning values to each ISBN. VCR manual from 1992? Probably not high value. Programming book from 1992? Higher. Scifi/fantasy magazine? Medium. Availability would then factor into it as well. When I say they might be buying up extras of some books to take them out of the market (or delay the competition getting them for a model version or two) I don’t mean a human is making a decision, or that they’re spending huge amounts of money on it. When I say I wouldn’t be surprised they’re doing it, I just mean a system that is opportunistic when the cost to buy, say, 5 instead of 1 rare books isn’t meaningful.

  7. “So a few of these books can be physically rare in the sense of being uncommon, but if they are next-to-valueless, it’s fairly obvious nobody has particularly cared about their survival up to now.”

    For someone with a history in publishing, and an article that is trying to de-sensationalize some of the more alarmist press around the behaviour of the major AI companies, this sentence is a disservice to most readers. A big point:

    ISBNs are not an indication a book was mass-manufactured. Anyone can apply for one, and they’re common enough is niche markets. From my own bookshelf, I can spot at least three categories of “has ISBN, is still reasonably considered rare and valuable”:

    The most obvious: A first edition. In this case, of Atwood’s Handmaid’s Tale. The Canadian printing, so the first of a few first editions. A collectible, with continuing interest to people who collect such things.
    A street photography book; “10 Years of In-Public”. A niche and high-quality volume, including a number of prints. Printed once for a specific event, never again. A number of photography and art books fall into this broad genre, including most of Magnum’s publications.
    Probably most relevant here: Making Kodak Film, a self-published booklet on the technical history of Kodak’s film emulsions, coating techniques, etc. written by one of their lead engineers after a decades-long career. Rare in the sense of irreplaceable knowledge, but too niche to be of interest to a mass audience.

    So, “has an ISBN so can’t actually be rare or worth preserving” strikes me as true in exactly the same way as the argument it’s trying to oppose: exactly true enough that anyone predisposed to agree with it gets to feel better about their opinion, but ignores the actual fact and nuance of the situation.

    And conspicuously absent from the article: how is this different from when Google did it, for Google Books? Because it is, in a couple of key ways:

    Google managed to figure out how to scan copies non-destructively when warranted. It’s slower and needs more labour, but the option is there and the cost difference is negligible in “frontier AI lab” terms.
    Google’s scanning let people find and access the work and knowledge the books contained. AI training is doing roughly the opposite; the books are trained into a model’s weights, and to the extent that they’re actually preserved and accessible if you come up with the right prompt, well, what do you then check it against?

    Simplifying the behaviour of these companies to “they’re not destroying medieval manuscripts,” “you wouldn’t believe how many Dan Brown novels get pulped” and especially “it’s easier to scan a pile of pages than a bound book” is a weirdly lazy position to dress up as insider insight.

    1. ” A first edition. In this case, of Atwood’s Handmaid’s Tale. The Canadian printing, so the first of a few first editions. A collectible, with continuing interest to people who collect such things.”

      They’re paying commodity prices, they aren’t getting that type of rare book. But, unless you can point out something fundamentally different between a first and X printing that matters, I’m still not going to care. The rare books they are getting are the first/seconding printing of books nobody noticed in the first place. Maybe there’s a granual of value that’ll be destroyed, but more likely they’ll be destroying a questionable treatise on the mold between your toes written by an age addled, drugged out professor with tenure.

      More valuable books are destroyed in an afternoon for other reasons than in these AI processing facilities. And ironically, being integrated into an LLM will probably give some of these more influence on future society than they would ever have by dustying up a shelf in your mothers basement. (Let’s be honest, it isn’t your basement yet.)

    1. The BNF has a website called Gallica with over 5 million books digitized, from antiquity to the 1950s. One can easily loose themselves in the beauty of medieval and renaissance illuminations. One of my favorite scans is their of Giovani Battista’s architectural drawings, with many more than the sole prisons.

    2. Not really how it works – some things that everyone notices are become fewer gain in value as the volume of them decreases, which usually requires a gradual decline of something that was already valuable and rarer so people are paying attention.

      But others nobody noticed it was the last copy and it didn’t seem valuable at the time – for instance consider the steam railway locomotive, here in the UK they became almost valueless as they all went to scrap practically at once when the railways went diesel/electric. If a few steam enthusiasts that would become the heritage railways didn’t buy and restore some of these wrecks there could easily have become none left of that model locomotive, maybe even none left of that era of locomotive now!

      1. But others nobody noticed it was the last copy and it didn’t seem valuable at the time

        Or the many missing episodes of “Doctor Who” and other such shows.

  8. I suspect that the AI companies where trying to keep this low key, because they knew people would be upset about it.

    I also suspect that most people learned about the story through a curated news source in which an AI selected the story and sensationalized the headline/summary to maximize the likelihood someone would read it. How ironic.

    Lastly, according to a helpful AI search agent, the most popular genres of books published since the 70s are romance novels, followed by crime/mystery stories. Great. Get your questions answered by a horny AI trying to commit a crime without getting caught.

  9. In re-reading the comments, it is unclear to me if the outrage for scan-n-destroy is because of the lost of accessibility to a limited physical copy or if the hostility is directed toward the AI industry who will (supposedly) restrict access to the scan until the contents are declared public domain?

    Neither did the categorization of the scan seem to matter; that is, works of ‘fact’ or of ‘fiction’.

    Facts we know are often period-centric and are often proved false in the future which makes them history.

    Works of fiction have historical value perhaps but it seems to me to be of minimum value to an AI unless a query was directed specifically towards the specific book or author.
    Example for Gemini:
    “Does the character Gollum stay consistent throughout the books series Lord of the Rings?”
    Response:
    Gollum changes significantly in his personality and motivations throughout The Lord of the Rings books.

    I am neutral on AI book scanning. But, to extend the profit-driven paradigm from written word to works of art, would there be outrage if Disney (substitute your fav monster company) had purchased at a recent Sotheby’s auction* the Klimt portrait and then scanned art, destroyed the original, and set-up an Internet accessible pay-to-view portal?

    http://www.sothebys.com/en/articles/the-new-york-sales-november-2025-breuer-results

    1. What makes you think that they will suddenly share them with humanity the moment that they are legally able to? These are the same folks pirating music and claiming that it’s fine because the originals aren’t accessible in the data.

        1. I assume they have a perfect scan, but only ‘train’ on the OCR output. It would be a shame to delete the scans after they are made, and given how data hungry they are, they would not.

          In the end, a much better use of the tech is to allow searches through books. This doesn’t need to be an LLM, though. It can just be a semantic search. This is quite similar to how Python (pandas, matplotlib) is used by an LLM for analyzing and plotting data, which is just not a thing an LLM can do.

  10. They started off downloading digital copies (illegally) and no one was happy about that. So now they scan them, and no one is happy about that. I get that some people will probably never be happy about anything AI. There were people that were never happy with calculators. Hell, there were people who were never happy with BOOKS!

    Life goes on, those people die, the future never stops approaching.

    1. Good points, really.
      On other hand, who decides what’s important and not?
      Books are written thoughts, people’s legacy. Our all legacy, even, maybe.
      If the books vanish, so will be the memory of these people.
      There’s a saying that goes like this: No one is really gone as long as someone still remembers him/her.
      By destructing the last copies of an “uninmportant” work, their authors are also being erased from our collective memory.
      So AI firms do much more than killing off books, they’re killing people post-mortem. Philosophically spoken.

    2. The future never stops approaching, but can you see what it is? I don’t know if “more tech, forever” is the answer to that.

  11. “‘Look at a book. A book is the right size to be a book. They’re solar-powered. If you drop them, they keep on being a book. You can find your place in microseconds. Books are really good at being books and no matter what happens books will survive.”—Douglas Adams

  12. The arguments put forth are weak at best:
    1. I’m not worried, so you shouldn’t be either.
    2. It’s okay because it saves the AI companies money.
    3. It’s okay because it lets the AI companies skirt copyright law, maybe.
    4. It’s okay because the books have ISBN numbers, so maybe they are not rare.
    5. It’s okay because whatabout all the other books being destroyed.

    This is a race between the AI companies to grab as much original material as they can for training their AIs. My guess is that an AI agent was tasked with getting a copy of every book. Books with ISBNs were the low hanging fruit. But it won’t stop there. AI agents will be searching online inventories of antique booksellers for new sources of data.

    It feels like it’s all on autopilot. The questions is where is the oversight? Who can say “Wait, we shouldn’t cut up this book because it’s value as a cultural artifact is greater than it’s value as training data.”?

    1. What are your arguments for no it’s not ok ? Except emotion.
      ISBN books = recent books = copyright = legally bind to destroy them.

      1. ISBN books = recent books = copyright = legally bind to destroy them.

        I’d say the moral imperative to preserve them outweighs the legal imperative to destroy them.

    2. This is what I would say as well. This is an apologetic article, sidestepping the problem at hand.
      And perhaps another thing: books are not just the words on the page, ask any antique book seller.
      I seen digitizing of antique books, they use a wedge shaped scanner and air puffs to flip pages. The problem isn’t digitizing of books, the problem is the sheer hubris of these companies, and their complete lack of respect for anything but themselves.

  13. What was the movie, short circuit 2? The robot just bursts into the bookstore and starts scanning books like a freaking lunatic. “Innnpuuuttt!”

  14. 75% of countries in the world have some kind of what is called “dépôt légal” in France = one sample of every book should be send to some institution for conservation.
    In France, this is mandatory for every printed document (book, flyer, advertising catalog…).
    So the BNF is full of rare (and probably unique) documents like the 84′ Christmas catalog of toys of Carrefour (our local Wallmart)

  15. This article got me thinking about (but sadly didn’t actually go into) what categories of books are both rare and not valuable. After some consideration I came to the terrifying conclusion that self-help books that didn’t help and Z-list celebrity biographies are both strong contenders, maybe we should be more worried about the influence of these books on the output of LLMs than their destruction.

    1. I have a lot of old (1920s – 1950s) engineering books. They are certainly hard to come by, and often quite interesting, but they have no real market value. Selling them is a total pain as the cost of shipping often exceeds the asking price; often this means that the book simply doesn’t sell, and eventually I chuck it in the recycling. Truth is, if someone is scanning these books and ingesting the knowledge before recycling them, they are probably one up on me.

    2. That is a very valid point. One of AI’s pitfalls is its lack of ability to weigh the credibility of a source. I recently posed a question in the comments of a different article, and after promptly being responded to by a naysayer, decided to research the question myself. I had a good laugh when Google’s AI cited random naysayer’s comment from 5 minutes prior as though it were a definitive source.

      1. Yes, this whole thing is so frightening because either way it has a bad outcome with big impact.
        If you train Big AI solely on value-less literature, the AI suggestions will be equally value-less, but wide-spread, which is obviously bad. The quality will somewhat go up if you fail to sort out the books with valuable knowledge (and it is probably cheaper not to have an expert categorize each book beforehand), so the system is promoting the undercover destruction of high quality literature with low volume. Even if you manage to train solely on high quality, you lose the credibility and you have to rediscover any line of verification all by yourself, which is much more work than finding and learning from a good book.

          1. that’s the big problem: how to decide what to keep? How to measure quality without missing something out of scope?
            For me “low quality” in this context means “I could have gained more valid information, if I had not used AI but some/any other means of information gathering”.

  16. Maybe one of the coolest hacks I’ve ever seen, though some time ago.

    Google Tech Talk.
    Usually these were like at the most 10,000 or less views on youtube. This one was already at 100,000 views when I spotted it, so knew it was going to be something.
    Basically, took a regular vacuum cleaner and turned it into a really good book scanner.

    https://www.youtube.com/watch?v=4JuoOaL11bw

  17. This is extra information that’s good to have, even though I’m not quite in agreement that there’s no cause for concern with regard to how AI companies are handling books.

    What I find a little weird is the copyright thing. So if you scan a book and especially if you then extract the text with OCR, that counts as creating a copy. That’s fair, I think. But from there I feel like it gets murkier because of how computers work. Does transferring the book from one drive to another counts as making a copy? Probably, right? But what about displaying the text on a screen? That requires duplicating the data in memory and then the frame buffer, right? Are those copies? And if you digitize the book solely for training an LLM, the resulting model may or may not fully contain the data. Depends on the size of the model and the total amount of training data, right? Is that a copy? And if it is, wouldn’t that imply that you copy a book by reading it?

  18. I’m admittedly not well versed in the topic, but the two questions that come to mind are 1. How many companies are doing this, and 2. How many copies of a given book do they need? In other words, what’s the scope and scale of this? If it’s a small handful of big tech companies and they each only need one copy of a given book, it’s hard to imagine them making a significant dent in the supply, whereas if it’s hundreds of companies or they’re chewing up several copies each, that’s a vastly different story.

  19. A lot of people here are really confused what a “rare book” is.

    I have a rare book about historical recreation of Cornetti (renaissance instrument). Single publish by a museum in 2012. They have no more copies. Can’t buy it anywhere, not even scalpers. No print on demand.

    Destroying to digitize without sharing completely online is an utter disgrace. What’s worse is, when trained in an LLM, its lossy-compressed. We can never get the plaintext back out.

    Its really a gross complaint towards capitalism AND copyright both.

  20. There are non-destructive ways of scanning books.

    Simply force them to do it that way. That ways society might have some benifit. Digitizing actually rare tomes is something a lot of museum would like help with. In that way it could server industry and society. Making sure people can again read the rare books that otherwise none may handle.

  21. Substitute ‘painting’ for ‘book’ and consider rare paintings. The ‘Mona Lisa’ is a rare painting and all of the companies can afford to buy it and scan it. Would then shredding it be ok? Would not the correct thing to do be to ‘return’ it? Why are the books being destroyed? They have been read. They could sell them back. But that would then tell others that knowledge had been acquired.

  22. “A few weeks ago, I wanted to re-watch The Dictator, with Sacha Baron Cohen. …
    For some reason, someone seems to have decided to prune about 20 minutes of content from this film….
    Then, I noticed a similar trend with lots of other comedies. A few seconds here, a few seconds there, the catalog on streaming services differs from hard-copy versions that I know
    To what end, one might wonder? Add to that the arbitrary decisions by media giants to discontinue parts of their digital catalogs at whim, decisions to restrict or even remove physical media, and you start wondering…

    “…Printed soft and hard copies of books are immutable, and what they carry is like the cosmic background radiation, a precious view unto the past, how things were, good and bad and stupid and brilliant, all of it. Books also don’t need a battery charge, and they work fine even after 40 or 60 years…”

    Cherish your physical books and media”
    Updated: August 28, 2026 | Category: Life wisdom
    Dedoimedo
    https://www.dedoimedo.com/

Leave a Reply

Please be kind and respectful to help make the comments section excellent. (Comment Policy)

This site uses Akismet to reduce spam. Learn how your comment data is processed.