2 days ago https://news.ycombinator.com/item?id=49310725
21 days ago https://news.ycombinator.com/item?id=49068738
I'm not really aware of many benevolent acts Amazon has taken in the past decade. Are you?
* They're only buying a single copy of that book as they only need one to scan
* If the book was public domain (or should be), then there should be an effort to "democratize" that data into a public commons of intellectual property?
Would that an enhancement to the Library of Congress or such?Why do you think for even a moment that they would buy only one copy, rather than every available copy of something rare and difficult to get? Anyone doing this stands to gain from destroying the only copies of something that has scarcity, to stop competitors from ever getting access to it. That's a serious concern.
Why would they do that? What is this 'should' to a company that would do this in the first place? Are you not suggesting the opposite of what they aim to produce?
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
Rare invokes images of limited edition runs of well loved books, when in reality it's probably extremely outdated software guides, how-tos, technical manuals, etc.
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
In more liberal sense, because you might be destroying something unique that later generations might actually like to see, the life rule that states "Don't be a dick" probably applies trumps even first sale doctrine.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
If AI-scanners are somehow bid-sniping bona fide collectors and wealthy aficionados, there may be cause for concern. But that is most certainly not happening here.
Just another reminder that piracy remains the absolute best archival strategy we have.
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
It is conspicuous that, aside from established booksellers, and perhaps the LFL owners, many of the people crying out and moaning over book-destruction seem unwilling or unable to actually take those books in and store them. That is the point, that these books are unworthy of taking up storage space, which is an ongoing cost, and maintenance, and therefore, it is fiscally responsible to sell them, scan them, and destroy them. If you want to dedicate a room full of shelves to books that nobody will ever read, then go ahead and buy out your local bookseller! They will not stop you or cancel your order! Tell them you are saving those innocent books from the Big AI Bugaboo! They'll grovel at your feet!
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
So books that would probably have ended up as trash. These AI training facilities are actually doing these book a service. Not only they are probably going to keep the scans safe (for future training), but having the book end up in an AI model may be the only way it is going to have any use at all.
Let's say for instance that the book in question is about woodworking, and it is not great, lots of mistakes and inaccuracies, unoriginal content, etc... except for a single thing, maybe a trick for making a certain measurement or something like that. Who would read such a crappy book for this single good trick he doesn't know is there, well, an computer will, computers process terabytes of crap without tiring and complaining, that's what they are for, and with a well designed LLM, that one trick may resurface, waiting for someone to ask about that specific measurement.
The problem here is not that rare books end up in AI training facilities, it is that these AI facilities are owned by for-profit companies keeping the data to themselves. These books should go to public libraries instead, for everyone to access, it would be better if these books weren't destroyed in to process too. But the question becomes: why didn't public libraries didn't do that in the fist place? And maybe in a more respectful way. The AI companies would just have had to license the database to libraries, probably simpler and cheaper than having their own scanning facilities.
To me, this mess is a failure of the copyright system. One one hand, large scale book digitization projects intended to preserve and make the original text accessible get lawsuits by publishers, while AI training is "fair use". It means we have built a system that encourages destroying rather than preserving these books!
This is the legal loophole that allows them to do what they need to do, and the benefit of doing it outweighs the modicum of outrage this title will generate.
Edit. Because I see my statement confused all the HN experts:
Anthropic's version of this was Project Panama [1]. The destruction of the original allows them to keep a digital copy in the AI training dataset. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
Here's a bookseller describing some of the books, which he believes are going to middlemen who are playing a sort of arbitrage with the AI companies by buying obscure titles, reselling them, and tossing away anything that doesn't sell. https://charliebecker.substack.com/p/is-an-ai-company-buying...
I suspect that's because folks wouldn't react the same way if they realized it was computer manuals for Windows 3.1. Even more obviously, Amazon doesn't want to be paying a lot for these books, so it seems outlandish that they'd be buying valuable rarities, since the booksellers would know the worth of those copies.
Good reporting on this would have included sales prices, volumes, and titles.
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
(this reasoning was rejected by the court)
https://law.justia.com/cases/federal/appellate-courts/F3/118...
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
https://law.justia.com/cases/federal/district-courts/FSupp/5...
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
I just looked for one novel, which wasn't a fantastic book, but it is the first use of a pithy and fun phrase that is so ubiquitous that you'll probably read it a couple of times today. I argued with Claude, GPT and Gemini for ten minutes just now, even knowing the title of the book, to even prove the book exists. It took me years to find a copy originally and then I lost it in a move. I found one more copy today from a rare book seller, but it just sold (to Amazon?).
Is the book valuable? Not particularly, but I feel it's noteworthy and important. I don't want to name it either, because now I have some searches out and the next copy that pops up I'll scan and put on IA. There can only have been a few thousand copies originally published in 1947, it's only in hardcover. I know of a couple of other copies in private hands, so it's not zero copies, but it has to be single-digits.
I have one periodical issue that I know of only one other existing copy (Worthpoint only shows one copy ever sold in their database) and if you look on collector sites there is a blank because nobody even knows what the cover looks like. I can't explain it, since the publication routinely printed hundreds of thousands of copies of each issue, but here we are. Perhaps all the copies were withdrawn and pulped immediately after publication for some reason? It's in my scan pile, so I'll have it uploaded soon. Is it significant? Not hugely, but every other issue of this title has been scanned already, so it's scratching an itch to get this one done.
There's definitely rare stuff getting scanned and shredded. Someone in a comment above said it's not like Nazi book-burning since they were trying to destroy information. But it is like that if you consider there remain no other physical copies and all the electronic copies are locked up in a way that nobody can access except to trick an LLM to spit out a paraphrased copy from its training data.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
It sounds shocking if you're missing the context and only rely on the blogspam which is this TechCrunch article.
They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).
The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.
Anthropic's version of this was Project Panama [1]. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
We've also seen a website that collects old paper restaurant placemats from across the ages. Literally disposable, zero value items that were made to be discarded by the thousands.
And yet, when someone has an interesting idea that ties them together, makes us recollect or think about our past in some novel way, these useless uncommon things suddenly become quite interesting.
A book on sermons, o its own, from the 1920s is maybe not that interesting. A book of sermons selected from each decade? A collection that compares regional books of sermons? A compare and contrast of the 2020s and 1920s? I can imagine many interesting thesis where a book like that becomes interesting because of the context of other books that are juxtaposed to it.
A book on Detroit motorways and bus schedules from the 30s isn't interesting or valuable on its own. But when you contextualize it, suddenly it might be a way to understand our history, our path through development and redevelopment. Connecting our present moment to the past.
We got started when another ministry moved out, and sort of from zero. My pastor's clear instructions to me were: make sure everything we carry is doctrinally sound.
So we inherited several full collections of books in rapid succession. Some had even belonged to priests and religious. Those gave me a fascinating time, because I could basically rubber-stamp every title that a priest had in his personal collection. But slowly the balance began to tip into rather esoteric volumes that normal laypeople couldn't really use. Literally books full of sermons and other arcane subjects!
I was tasked with discarding/recycling all the rejects. There were tons of rejects, believe me! So with every session when we had boxes full of donation, it was imperative to cull the bad stuff very fast, shelve the rest, and then find somewhere to dump the trash. The manager was encouraging me to recycle, or at least not tip them all into the Dumpster, but it turned out to be a logistical nightmare to find anyplace that would recycle books like that. Having no vehicle, I had to continually figure out ways to cart around heavy loads, just to get them out of church and into the trash somewhere. That was the worst part of my job.
Now it was clear that books were not a very hip or current medium, but there were plenty of elderly parishioners who did appreciate the resource and did compliment my work, but our church was not free of prejudice or judgementalism, and let me just say, there was an angel or entity whose purpose was only to jumble all the books while I wasn't watching, and leave a deliberately unorganized mess for me to confront every week. This made the task distinctly Sisyphean, in addition to the need to constantly discard rejected books.
I finally threw in the towel when large boxes of Spanish-language books were donated; there was no way at all for me to vouch or determine their orthodoxy, and the shelves were full anyways, and I was just tired of propping up a legacy ministry anyway. But it really drove home my opinions about books, hoarders, and that is why I have no troubles with the way books are currently being treated.
I didn't realize that the list of books was published somewhere, so that you know with confidence what they're adding to the collection.
How do I browse the list you're using for tracking this?
Of course I could be wrong and they are destroying 14th century monastic scrolls or out of print Mills & Boon editions.
And that is a reasonable assumption?
Personally when seeking out rare books with few copies in existance and extremely scarce availability, I have often found listings that have lasted for quite some time. Not all rare books go to auction. You might be thinking only of some extreme of notable works and not a wider spectrum of desired but scarce publications that does indeed exist contrary to your confident assertion.
the part where the article avoids mentioning what book this was is a tell
The verbatim content are the words, not the paper.
Books are lost all the time because the last book ended up in a landfill. But if an AI lab digitizes it, now it's stored in an extremely redundant storage lake in a datacenter and the company has huge incentives to make sure they don't ever lose that data.
You're literally assuming the conclusion. The exact topic under contention is whether they are "destroying human cultural heritage".
"They're allowed to be a hypocrite" doesn't mean they aren't a hypocrite.
While I wish there was a repository of every book that was already digitized (it pains me this is the best solution), there isn't one and so I think this is not a real problem.
It'd be a different story if they had furnaces that ran only on rare books that they had to continually feed books to but that's not what's happening here. And that 1 destroyed copy will "live on" in a way that it otherwise might not.
Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
It seems a little short-sighted to destroy an artifact to get the text.
Perhaps the genetic material that remains in books from the people who handled them, or the pollen from plants in the environment that the book existed in will have value in the future, but we won't know what we lost because some people foolishly destroyed it in a bizarre quest to make AGI that the creators argue could potentially destroy humanity.
The more and more I read about these kinds of people the more I'm starting to realize that they're in the "here for a good time not a long time" group of people and those are the last people you want making long-term decisions.
The sad truth is the vast, vast majority of printed literature is neither interesting nor useful. People are not dumping off stacks of Umberto Eco. We frankly have to toss a lot of awful cookbooks, self-help books, trashy mass-market "novels", and sketchy religious works. As it is, even the stuff that makes it to the library is not very impressive.
It's all interesting for something, even if it's just a meta analysis of culture during a certain period or what kind of trashy romance novels were popular in 198X. At least in my view.
The latter seems inefficient, so my first assumption would be that they would avoid that. But while I suspect they'd check if they know the book before scanning, I could imagine them not caring that much before buying them and just focus on volume.
That's the point I come to.
Books are not original manuscripts. Even in low volume cases, they are usually printed hundreds of times. (And usually low volume works aren't all that great...hence the low demand.)
That is a fraction of a percent for books that at some level weren't all that wanted.
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
Citations? Also what exactly are these 'normal means'? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.
Normal means is throwing it in the recycle bin. Especially for stuff like a 1982 John Deere manual. I've never donated an old appliance's manual to the library. Have you?
Amongst my friends, I'm one of the rare folks who donates books to the library. Most people just trash them. And I know the library only wants them to try to sell them in their book sales (or online) so they can get money. Almost nothing one donates to a library actually ends up on the library shelves.
Not even the title of one of those rare books?
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
So far from how I see how powerful organizations work, I'm not as certain as I would like to be.
People say this all the time, but so far nothing has convinced me it's true.
LLM development has more or less plateaued, and the current boundaries are very real - energy, resources, capital.
At this point we're talking about marginal improvements against the same asymptotes of all technological innovations.
You could easily write an article about how Goodwill and the Salvation Army dump millions of "rare" books that no one would ever buy for $0.50. And thrift stores will dump multiple copies of said books. AI companies would only ever need one!
The only thing I disagree with is AI companies feeling entitled to freely use every piece of copywritten work without restriction, but that applies just as much to web scraping, pirated books, etc.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
(This is why I will never be a billionaire)
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.
I have a few "rare books" and have read many, you'd be suprised at what is publicly available on google books since like ~2010ish.
The first was destructive. This was for mainstream books currently being published so they had no value. It's (I believe) where you cut off the spine and scan the pages.
For rarer books, there was a non-destructive process. Basically the book was opened to each page and scanned. This was slower but didn't destroy the book.
I don't understand why these companies haven't just licensed the scans Google has already done. Why is each company doing this rather than just scanning the books once and sharing the scans?
Because for the vast majority of the books, Google doesn't have the legal right to license those scans. It would be legal if the books were out of copyright, but despite the connotation of "rare books" those generally aren't the books we're talking about. Further, in the cases where Google didn't destroy the original of the book they scanned, their scanned copy may be considered infringing under the new standard, so Google doesn't want call undue attention to what they have.
https://nitter.net/internetarchive/status/135809098218971955...
6 Feb 2021
At the Internet Archive, this is how we digitize a book.
We never destroy a book by cutting off its binding. Instead, we digitize it the hard way--one page at a time.
Wonder the additional cost to add automated page turning. More than Bezos can afford, impoverished chap he is.It was a distinction between library books and others. They partnered with libraries, and obviously, libraries didn't want their books destroyed, so they devised a non-destructive book scanner. Some of the books they scanned were indeed rare, but rare or not, you don't destroy books you borrowed from a library!
(Also the person who coined Singularity, though Ray Kurzweil really wanted everyone to think it was his idea.)
Because it's not just one company doing this, it's dozens if not hundreds of them, all with huge budgets and all competing over the same dwindling supply of older books. What those talking about library disposals and similar measures are missing is the sheer scale that destructive book scanning for AI ingestion is operating with. It's unprecedented.
It's not impossible that in a few years the number of older physical books for sale anywhere will plummet to almost nothing and that in many cases that'll include the last copies anywhere of particular titles. Second-hand book stores will close, and a ton of niche knowledge might be lost forever. It'll also all have been done in relative secret, without any transparency or public record kept of what was lost.
Worse, it's something that can only be done once. If our societies don't do something about this now there isn't going to be a chance for a do-over. Once a physical book is destroyed it's gone forever, and we'll be lucky if there's a digital copy left. But even then digital copies don't provide the same level of forensic verifiability that physical copies do. Future researchers looking for material that might've been contained in books like these will be out of luck.
If our societies do nothing to stop this, we'll all be poorer for it.
What worries me about the trends is inevitable sanitization of content or straight out falsification.
That sounds really speculative, and not inevitable at all.
I think you're just trying to invent things to be worried about because you don't like AI and don't trust AI companies.
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
The recent trial court decision that keeps being pointed to to support that:
(1) Found that for training AI, digitizing and copying works was fair use, period, with no requirement to destroy.
(2) For creating a centralized digital library for general use, digitizing works while destroying the hardcopy was fair use even with no intent to use them for training AI.
“And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.”
I’m not a lawyer, so it’s possible I’m missing something. But it seems to me like the ruling here implies fair use only holds so long as no net new copies, digital or otherwise, are created. If that’s the case, then it necessitates the destruction of the original.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
I think the common consensus is that stackoverflow killed stackoverflow, quite a few years before LLMs became entrenched.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
https://www.authorsalliance.org/2025/09/07/the-anthropic-set...
But they won't share them because they don't want their competitors to have the data.
And if the books are not in the public domain, then they should not be allowed to train their AI models with the material without some kind of license or agreement with the owner of the copyright.
They don't care about the R.L. Stein Goosebumps series or Louis L'Amour pulp fiction. They want the stuff that isn't digitized and they're willing to destroy original artifact to do so with no guarantee that even the digitized text will ever be made publicly available.
It's foolish and selfish.
I think the more important part is to define what "rare" actually means and save/digitize set copies of those books that are actually at risk of being lost completely (rather than letting them rot away somewhere or get bought up to be privately destroyed). I suspect many of them are just not at all interesting enough to justify the expensive though.
As a general rule preserving historical artifacts for continued future analysis and appreciation by people who have yet to be born is a noble cause that's considered worthy in and of itself without the need to justify it. We're talking about the richest group of people who have ever existed with the means to preserve them but decline to do so and instead act like the Taliban blowing up statues of Buddha.
I think this is intuitive to most people. If a prerequisite for the birth of AGI was that a humanoid robot had to sit down in the Louvre and eat every single painting there with a knife and fork while occasionally stopping to wipe away historical detritus with the Mona Lisa that it's wearing as a bib people would by and large express a visceral and justifiable outrage.
The people who are doing this know how this looks so they're trying to do it all behind closed doors, through cut-outs and intermediaries.
My neighbor self-published a book, printed I think 100 copies at his own expense. It's literally a rare book. I can't imagine he nor anyone would care if an AI company bought a copy, no matter what they did with it.
Maybe these "rare books" are first edition Mark Twains, and. maybe they're unwanted books that would otherwise have gone to be pulped. The distinction is important and by not giving any evidence or even a qualitative claim about the types of books, 404 media is being pretty weak here.
And beyond just benefit to society, don't forget how subjective the decision about what constitutes a "unique artifact" is. That random product manual from 1966 becomes a near-priceless relic to me if it lets me repair and use the sewing machine handed down from my grandmother. Not necessarily the ink printed on paper, but the information contained within.
The above is based on a true story involving archive.org's Manual Library. The idea that even such esoteric and forgotten information would get hoovered up into the walled garden of Amazon's AI and then destroyed in the outside world, such that I have to go pay them to access even a facsimile of it, is frankly disgusting.
On the other hand, now that the manual is in the AI, you can just ask the AI how to repair the sewing machine.
The information contained within has not been lost, and in fact has become much more accessible.
Like so often in its usage, the word "just" is carrying immense weight in this sentence. Please see "hoovered up into the walled garden of Amazon's AI [...] such that I have to go pay them to access even a facsimile of it"
> has become much more accessible
This accessibility rests on a number of assumptions, not the least of which is Amazon's (of all companies) charitable good graces in offering access to their AI at an affordable price. It also assumes that the model can accurately regurgitate the text without hallucinating about other, similar machines, and that it can faithfully recreate any diagrams.
It would be great, but it's a bit off topic here. The point the parent is making is that most of these works are going to be trashed regardless of Amazon's behavior. If Amazon (or any other company) is digitizing it, they're at least preserving it in some form. The alternative may well be that many of these books never get preserved.
Now would I prefer the government or an entity like The Internet Archive do this? Sure.
However, a title can be cross-referenced against purchase histories in probably under a minute.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
https://www.gadgetreview.com/we-dont-want-it-to-be-known-ins...
https://www.irishtimes.com/world/europe/2026/08/10/a-mysteri...
https://www.bbc.co.uk/news/articles/cp3rprx2wl4o
https://dallasexpress.com/national/the-vanishing-page-ai-fir...
https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phis...
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
Real humans sharing their unique knowledge, packaged in a book.
https://cdlib.org/west/ https://papr.crl.edu https://eastlibraries.org
Actually I expect that a copy of most books probably is still available in a copyright library but it might not be easy to access.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
I would be also interested in participating in your costs.
Materials like that may be very valuable to software / tech / HCI archeologists soon.
I can't get the tools or local know-how to straighten my scythe blade in a country where every cottage had a scythe less than 100 years ago, with the last scythe-native generation rapidly dying out.
Fast forward a civilizational collapse and that M$ Office 97 for Dummies might be as groundbreaking as a Guide to Using Roman Concrete
Thank you.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
If people are really concerned about bison, they should lobby gamekeepers to release photos of them.
https://en.wikipedia.org/wiki/American_bison#/media/File:Bis...
I think this is a bit of a myopic take. It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
I'm with you on the "rare" part. If people are thinking about 70+ year old documents or ancient manuscripts, I doubt that's what AI is being trained on and is being destroyed, but it's reasonable that people find _that_ idea distasteful.
You can say it's manufactured but if these companies ignore this criticism, it's just another way AI companies are committed to losing the public.
Yes, and when you buy them with that designation, you know what you're getting into. The problem occurs when you already own it, and some group is trying to get it labeled as historic, which will add to your burden and limit what you can do with it.
When I was in a small town, this was actually weaponized. A hotel owner was trying to get another hotel categorized as "historic" and had rallied a lot of people behind his cause. He had a case - the hotel did have some claim to being the "first" in some category or other. But really, he was doing it because it was a competitor. The "historic" hotel owner had to spend a lot of money to fight the cause, because being labeled historic would prevent him from performing various upgrades, making the hotel less attractive to customers (he was already not getting many customers).
Sample:
> Thou shalt commit adultery.
They raise the price and print more copies?
It's the scale that matters as first, and secondly, most people don't shred their books after reading them once or twice. This is just beyond words.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.
The rare qualifier is used precisely because no reasonable person thinks this is stealing.
Nothing good comes from hoarding whether it being toilet paper, money or knowledge.
No reasonable person thinks buying 1 of something is "hoarding".
But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.
https://software.annas-archive.gl/AnnaArchivist/annas-archiv...
Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.
A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.
That's literally the only reason they are trashing them. It's a legal requirement.
Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?
For instance there were only limited production runs for a set of books on South Africa's participation in the Second World War, and getting hold of them is increasingly difficult. But they have ISBN numbers, so they fall within the group of books being collected and destroyed by AI companies, and as a result might disappear altogether, along with the knowledge inside them.
That's not happening though.
Because none are.
In the UK weeding is at the discretion of the Branch Librarian. Books I checked out 15 years ago have now vanished from the system.
This is a grave problem.
> you will (if the model is good) be able to get the knowledge out of it
That doesn't help with books that aren't factual in nature. "Getting the knowledge" out of a novel makes no sense.
Simple example:
Actual book from 1732 (rare, original, older than your country): https://blackwells.co.uk/bookshop/product/The-Compleat-City-... -- yours for £1,258.00
A scan of the same edition of that book: https://archive.org/details/bim_eighteenth-century_the-compl... -- free. Nobody's paying any money for the digital copy.
I don't think you understand the book itself is a collectible object with rarity and value, regardless of the information it contains.
Digital copies can be altered at whim, as the only means to provide any sort of tracking or verifiability is through another external software system that itself has to be trusted.
> What are the other people who own these books doing besides having them sit on a shelf?
Not destroying them. By continuing to exist, it keeps the possibility that the books will eventually go to someone else. Maybe even a library that specializes in rare books.
Good news, there's hundreds or thousands of copies that continue to exist. Anthropic, Google et al only need a single copy.
And shockingly, for all this talk of rare books, we don't see the other owners of said books stepping up and offering to have these books scanned non-destructively. I won't hold my breath for them to do so in the short or long term.
FURTHERMORE, the books aren't being scanned for no reason, they become part of new model training data. Meaning that whatever insights these books contain might be available to users in the future. Again, more valuable than all the copies rotting away on someone's bookshelf that might be scanned and made available some day.
I would check the settlement db just to make sure: https://secure.anthropiccopyrightsettlement.com/lookup
You may entitled to a share assuming the publisher registered with the class action and there is a copyright registration in your name.
I mean, Anthropic isn't going to fight it because it lets them do the thing they want to do, so I can see how this never gets beyond the court that allows them to do the thing they want to do.
But would this argument would have flown in the past?
It wasn't even attempted in Sony v Universal. Or any copyright suit up until this point. That doesn't smell funny to you?
The flip side is that Alsup (the judge who wrote the opinion) is probably the smartest district court judge we have when it comes to technology, and one of the people I'd trust most to come up good decision.
He's a treasure, and my instinct is that he got it right: https://en.wikipedia.org/wiki/William_Alsup
AI companies are not doing the same.
Far more people will use the model than would ever have read the book.
It's not at all the same thing as still having the books themselves available somewhere.
People may not even know the book exists or that it contains the information they want. Obtaining access to a copy may be very difficult, even if it is not particularly rare. The AI is much more accessible.
You're defending the indefensible here.
They can't legally publish the digital copies, and you know it. Training is fair use but direct copies are not.
You should be pushing for better copyright law not Amazon for buying up worthless books.
You might be right that they can't do that now. The simple path forward is to have these companies simply make a public commitment to publish the data when the copyright expires. I also think there is a space to be carved out, probably through regulation, to ensure that there is a clear path for these scans to enter the public domain, at the very least.
This is not just important for these specific books, I have written on other platforms and in other spaces about the importance of media companies and those who benefit from strong copyright laws to protect and generate profits and revenues to repay the public for the cost of that enforcement over time by ensuring that at the appropriate time, those works fully enter the public domain. That could mean a restructuring of the Library of Congress in the United States to become a modern Library of Alexandria to host data, and shifting to a registered copyright model where to gain the protections of the court, you need to upload/submit your copyrighted works for storage and eventual release. I doubt it would ever happen there because of the amount of money invested in tying up IP in the United States, but perhaps a more amenable location like the EU could help with that.
However it might work, part of the promise of the Internet was that information would be liberated, but we see every day how much information gets sent down the memory hole when businesses, sites or services shut down, or how regularly companies abuse IP related regulations to attempt to strangle competition. It would be expensive now, but it would create an incredibly valuable legacy of information for the future, and it will only get more expensive to build such a thing as time goes on.
If AMZN, et al. were burning books in good faith, they would be open about it, point to the party responsible for it, and provide a list of the burned books to allow the community to organize, scavenge and scan the endangered publications. Keeping everything secret is a proof of maliciousness.
Preserving books and providing easy access to the information in them is of utmost importance, this is certainly understood by the public and private entities who could do something positive about it - their actions in the opposite direction should be a wake up call, they're on the wrong side of this issue.
It would make more sense for government to accept digital copies for any book, and share the out-of-copyright ones. In the United States, we have a Library of Congress who could do it if a law were passed with funding.
Of course, because each company is happy to burn it all down to beat their competitors and being first to any content is most important, this will never happen.
Doing the scan within their own organization does not because it's within a single legal entity and there is no copyright issue involved.
Just imagine 20 years from now its hard to get books in print. AI companies can just change the history by altering their model's content.
I'm not to keen on corporations holding the world's entire print history in AI models.
Same answer? They raise the price and print more copies.
You can just print even 1 copy, just the cost per print will be higher.
In fact publishers getting some profit on the tail end of books is the best thing that could happen to authors and publishing industry.
How? Like, do you think a book published 100 years ago has a digital file sitting around a publishing house that might even exist anymore? Re-publishig of old books are often based on scans (the originals were literally pressed into paper by metal, not printed off hard drives), and if the books don't exist, they can't re-scan.
I own some antique Japanese books that have very few copies in existence. You couldn't get them re-printed if you want to if the physical book didn't exist anymore. For a few books, I've actually purchased incredibly high-quality scans that cost me hundreds of dollars because to produce the re-issues, people had to go to museums with high-powered cameras to photograph the pages under supervision of a curator. There's no digital file to print from the publisher. If those books were gone, there's no bringing them back.
You could overcome it through various means, including helping to set up a clearinghouse for bulk rights, implementing a data clean room approach that models could train on in situ, lobbying Congress for copyright law changes especially around orphan works, using Section 108 of the Copyright Act to set up a specific preservation vehicle like the HathiTrust, and other options.
I find it ridiculous that so many in this thread are acting as though AI companies are simply powerless to do anything but buy up and destroy these rare books.
Because people are irrational when it comes to this topic. They say stuff like “book burning” when they see books being destroyed. Doing it in relative secret kept the crazies away.
If you are passionate about this topic, lobby for more funding to store archives in the public good. Expecting private parties to do it for free because it gives you the ick to see books destroyed is not useful.
More books get destroyed each year during estate cleanouts than any AI companies could ever hope to accomplish. Most books donated to goodwill or other thrift shops go straight to the dumpster and might not even have a set of human eyes put on them at all. These places act as sin eaters for folks to leave their trash with.
They don't have to scan and destroy these books, it's an entirely voluntary choice.