Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.
Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.
At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.
There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.
In short, it's a mess.
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
These generally are not books people care about. The information contained therein was doomed.
Now the information has been digitally preserved and a digestion of the information will be made publicly available.
No, you have a translation.
>Even if, like the Odyssey, there are hundreds of wildly varying translations?
Precisely why translations are not considered equivalent to the original text.
>What about an abridged copy? What about the Sparknotes version?
An abridged copy is not a copy of the unabridged version.
>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?
No.
I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?
There's no technical reason why an LLM couldn't reproduce verbatim some of the training material. It's sort of a lossy statistical compression engine. Enough of the info will survive to the output in the original form. With the amount of data and the commercial nature it's hard to argue fair-use. But nobody tested this in court. I'm not even sure the US wants to ever test this. Why even attempt something that has a non-0 chance to sabotage your most promising industry/bubble in ages?
But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.
I'd rather a digital copy exist in someone's hands than a rotting physical copy.
It almost seems like you're suggesting that having Claude generate a paraphrased book is as good as having the original book but i don't think that could be your intention?
One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
https://genome.ch.bbc.co.uk/about
Historic TV guides are also the sort of strange ephemera that people collect. They ought to be digitized like newspapers and other magazines, but this was always the purview of libraries anyway.
https://en.wiktionary.org/wiki/wastebook
Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.
So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.
P.S. I'm not sure why you need to make fun of your own ignorance? Just look up the word you don't know and don't mention it?
And refusing to do this exercise just means that you behave as-if you put a really silly number on the value of human life, and probably not consistent between different parts of the project.
So I don't quite agree with these taboos in the absolute.
(I'm still against capital punishment on practical grounds.)
For books it's similar: if you taboo book destruction for the AI training folks, that doesn't rescue books from their ordinary pre-AI life cycle of getting destroyed all the time in the course of running a publisher or a library or a second-hand book store.
In fact, the AI craze is what's giving rare books _value_ and incentivises people to dig them up and preserve them. Or at least preserve them long enough to be scanned.
The scanning might destroy the physical copy of that book, but they save the contents. That's the whole point of scanning after all.
Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.
My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.
If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
They just don't want to pay what the copyright holders want to charge
I’m not aiming this at you directly by: ISBNs or STFU
Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.
People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.
It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.
BBC good enough for you?
https://www.bbc.com/news/articles/cp3rprx2wl4o
"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.
If these books are so rare and important, then why has no one cared until now to actually preserve them?
The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.
Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.
Why are AI companies forced to shred books?
This is just one example, but it has become unfortunately common across all social media platforms.
That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.
So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.
Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.
This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?
You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.
Source? You state that in a tone that implies you have verifiable knowlege of this. All the information I found says that the exact number, titles and authors are under NDA.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
The law is currently forcing these companies to destroy the books after scanning them.
Not "need to ingest"
Copyright holders are capitalizing on laws on the books just like Jeff Bezos companies buying their own copies to shred
So in the end it's really a Congress problem as usual
Nothing forces them to shred books, they do it because it's slightly cheaper that way.
https://en.wikipedia.org/wiki/Project_Panama
So I stand corrected: at least some don't do it because it's cheaper (than to buy a license, or simply forego some things), but because they're fucking evil, or so stupid that it effectively is the same as being extremely evil.
On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.
They are not forced to shread them by copyright.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
Who is going to claim a copyright violation?
Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?
Most developed countries have a 'legal deposit' system with a national archive/national library that requires publishers to send a copy of their works to them. They've existed in some form for centuries in some countries.
Example for the UK - British Library guidance: https://www.bl.uk/services/legal-deposit
I support Anna's Archive, by the way. Information wants to be free.
To watch the world cup I had to spin up a VM in Brazil to watch it with Portuguese narration because the free transmissions are region locked.
I would gladly pay 5 bucks for it if it was possible otherwise and avoid the hassle.
On Brazil the world cup was being transmitted on youtube. in NL only on traditional TV channels or Online for the same channels (all for free but in Dutch).
And literally as I write this I receive an email saying that my youtube premium was raised from 33 to 38 EURO. So there we have piracy getting juicier and juicier.
My concern is that less popular content is just erases, lost in mergers or lost in massive datacenters, never to be seen again.
I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.
If they we’re just kept where they were you can always say they’ll eventually be scanned or have that potential.
I do believe the issue is a bit overblown but the core of it sounds reasonable to me.
If the books are rare enough there's a shared cultural value that is being destroyed.
Just because they are for sale doesn't mean that no one else would have bought them. Destroying them obviously destroys them, which obviously destroys culture. After all, if the books is destroyed, it will not be read.
Think of it like this. If I buy an island, and there are bunch of people who live on it, maybe they're day laborers or whatever, is it right for me to expel them, if I don't want them? Can't that be straight up genocide, if they have some culture indigenous to the island?
The same is true for books. If you destroy culture, you destroy culture. There is no magic which triggered "but I owned the physical book" which means you didn't.
It doesn't matter what property relationship you have to a thing: whatever property relationship you may have to it, you still do what you do, i.e. you take whatever action with respect to it as you take with respect to it.
Another good example is food. If there's a shortage of it, and you still have a great deal, you may think "eh, what does it matter if I accidentally burn some, I save time through my carelessness" but if there are others who aren't getting any, you may be killing people, and may be despised for how you use "your property". The same is true here. Why should one not despise one who deprives others of rare texts, that may even be lost, and thus cause cultural destruction of some culture that may actually be rare, by treating the carelessly. That rare book on model building or whittling from the 70s may actually be important.
You scan books though and all the value they provide is still fully available in bit form, because books merely contain information. And of course being in bit form means they can be trivially copied, technically. But there's also the legality, which determines how many copies it's legally permissible to retain. The latter becomes the bottleneck for your preservation of "culture" because said culture could be made trivially available to anyone with an internet-connected device, if only rights holders weren't screaming foul (remember what happened with Internet Archive during/after Covid?[0]).
Ultimately there's no destruction involved when books are scanned, just a change of format which helps to ease management and for legal compliance.
[0] https://www.libraryjournal.com/story/internet-archive-loses-...
Which is also not illegal and within some bounds and exceptions, a protected right across the globe.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
This is what caused the problem in the first place. If people had unrestricted access to content then the world would be a better place. And works wouldn't be so rare that it's worth AI companies buying and destroying them to gain some edge, as well as remain in legal compliance.
A targeted modification to the laws could produce most of the benefits for a minor cost.
Like if a game isn't available for sale anywhere in my region, I can get it without breaking any laws. A book is out of print and I can't pay money for it digitally -> free game.
(Why "fair market value?": So that skeezy publishers don't have an online shop with one physical copy of every book they own for $1Trillion just to fulfill the law)
https://www.libraryjournal.com/story/internet-archive-loses-...
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
No. They also use lots of other methods to get training data.
There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.
At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.
Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up.
Edit; despite the above, looking at the court documents from the Anthropic case, this is pretty close to what they were arguing: “we are just transferring the physical form we purchased, therefore it is legal.”
I still dont think there is a requirement to destroy the book, but since there isn’t a reason to store the book and they can’t sell it, they probably just took the cheapest route. It might be worth an argument that they only purchased the right to use the digital copies while the physical copies exist, but I’m in over my head from a copyright standpint
I imagine since the law recently cost one of them truckloads of money for their violations of it?
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.
People write without any profit motive today. It's weird of the OP to think of writing in such a narrow space as commercialization.
Because there would not be human written books about the present. All books would be about the past. But literature is not stuck in time. Today writers talk about topics and feelings that writers of the last century might never know or experienced. Many people read books to better understand the today world (non-fiction) and to better understand their today feelings (fiction).
> People write for hundreds of reasons other than to make money.
Agree, but most of the writing that we have from the past still came with some form of financial incentives. Shakespeare didn't write all of the compositions just because he wanted to express himself. He was making money with theater performances. Many religious writing got patronage by the church. Dante Alighieri had a career as politician, Plato came from an aristocratic family. Writing was reserved to elites because education was expensive and people had to work for food.
Today we are lucky because education is accessible and printing is cheap.
> they created literature before copyright was a thing.
Copyright wasn't a thing because replicating content was hard. Try to manually copy a book...
Second, people write to be read. It takes huge amount of effort too. With no potential reward for it at all, they stop.
Third, we are social animals. If you dont see people reading, if you dont read yourself, you wont even think of writing.
there have been many times I've wanted to pay for an ebook but found that the only places that sell it apply DRM to it, so shadow library it is
Which exists to enrich Disney and other large corps, while they hide behind artists as a human shield.
>Artists need some form of reward.
Right, as do artists who use the work of other artists as their starting point. Copyright holders aren't bill and bob artist, they are massive corporate trolls throwing around the weight of almost 100 years of our cultural heritage, sucking the marrow from its bones.
>On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
The books are getting digitised into a permanent record of all our cultural heritage. It just sucks we don't have control over it. If only there was a way we could get them digitised AND control our cultural heritage. HMMMMMMM.
>Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write.
The incentive to write is being killed by slop groups like 20Booksto50K and Kindle which predate AI by at least a decade. AI just lets them work faster.
>We might end up with all of humanity’s books digitized and accessible for free
Excellent
>But there would be no human writers left.
Unlikely, but there would definitely be no Disneys or Conde Nasts left, which is a massively pro social outcome.
>In a world like that, what motivation would we still have to read?
In a world with all books digitised and accessible to read? A huge huge huge incentive. I already partake if books are too expensive where I am. It would take me the rest of my life to read all the books I already want to read. What kind of inane dribble is the idea that copyright makes it interesting to read? I havent even read all of Howard and he's in the public domain (in cool countries at least)
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
No data => No models => No competition.
> ChatGPT […] originally released on November 30, 2022
https://en.wikipedia.org/wiki/ChatGPT
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
Furthermore, they like that AI is bad. Because they think it's bad, and being right feels good.
> People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book.
Please pardon the tangent: that's what always bothered me about the Borg in Star Trek. Why do they need to assimilate whole species? I'm sure there are enough volunteers in the federation that would join the Borg collective. Even a handful should be enough.
So then they're gone :)
When the borg is threatened ("threatened") by a race, they choose to effectively end their way of life one way or the other. That is done by assimilation or by death.
If only a single outstanding member of a species becomes a threat, that's a strong signal for their overall potential; therefore everything needs to go, the sooner the better.
The crux of the issue is that it is mostly a PR problem. AI is amazing but being promoted, in the eyes of many, by the worst people imaginable. Very akin, and overlapping in many ways to the crypto crowd.
What if people like food more than AI? Have you considered that?
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
There are were polls about people being "more worried than hopeful" about satanic cults or alien abductions? With the youth being the most worried about satanic cults? Just like that knee-jerk "it's the bigger bubble", that makes zero sense.
As per the GDC 2026 State of the Industry poll, "52% said gen AI is bad for the industry, nearly double the 30% who held that view last year". But sure, everybody but HN, LinkedIn, and X bros are just luddites clutching pearls in their tiny bubble. They're the weak and stupid ones, and that is why the stupid shit said about them, day in and day out, isn't actually stupid. It all checks out.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.
I have found that often LLMs are familiar with the contents of SF books.
But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.
I'm sure they have some repository available somewhere. They can even sell the digital copies down the line if they're done with them (only once, of course).
Not in the EU, UK, or US. "AI" companies were already forced to settle their piracy cases, but often they get a free pass by law enforcement via regulatory capture.
The problem is a book author contracted publisher does not assign legal rights of duplication to a company/individual that buys a legitimate print. It can take over 70 years in some places to become public domain.
The core issue is "AI" firms have so much borrowed cash around, that getting a $1.5B fine for being a pirate is taken as a cost of doing business. The law is simply not equipped to handle this type of hyper-scaling criminal act. =3
Anthropic's crime wasn't stealing the contents of books and making a derivative work of it, but torrenting a shitload of books. Had they bought all the ebooks, I don't think the lawsuit would've gone anywhere.
That is not how copyright/trademark/contract laws treat similar works. Most LLM know about Disney Mickey Mouse, and LLM vector search space proximity results will gravitate more accurate reproductions of protected works regardless of granularity of data.
OpenAI simply canceled a popular service to avoid Disney wrath.
https://www.theglobeandmail.com/world/article-openai-sora-di...
Also, trying to escape directly ripping off notable famous people with nonunion talent:
https://www.youtube.com/watch?v=YhgYMH6n004
I would say the "AI" firms will keep buying time with all that borrowed cash. Everything that could be scraped has already been stolen, and thus the problem should begin to self-correct. The Shrek movie release market correction history correlation is funny, and a new film is due 2027 in July. =3
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
https://www.techbrew.com/stories/2026/01/28/anthropic-ai-boo...
-- Thos. Jefferson
A partial solution there is of course a larger budget and a "last copy" policy where the last copy of a text at least is stored away in deep storage against a future loan.
The local libraries also accept book donations for an annual fund-raising sale.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Biggest problems:
- scanning items that are bigger than the scanner platten so the start/end of every line is cut off.
- becoming an "editor": scanning only the pages you think are interesting and skipping intros, forewords, title pages, copyright pages etc
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.
Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:
(preview size, they sent a 500MB TIFF)
Now I can reassemble the issue and upload it.
I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.
I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.
More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...
It's that simple, these corps have once again broken the social contract, you must not reward them. OpenAI and Anthropic especially, both owned by schizo sociopathic elites. Just use Chinese open models on 3rd party providers or more ethical companies.
This is literally the only power you have outside of Luigi, you're not going to fix anything with a letter writing campaign. We are entering a fight for survival so you really need to step up your game and stop letting elites run over you.
Today it's just books and manipulating society, tomorrow they will track and punish your behaviour and the control will only get worse. These people are pure fucking evil and we need to start acting like it while we literally still have the freedom and privacy to organise.
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
And this is supposed to be concerning?