70 pointsby speckx7 hours ago22 comments
  • zirkonit6 hours ago
    Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

    I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.

    • szszrk6 hours ago
      My wife loves bulk book hauls. Those places that we frequent, work like a literal permanent discount warehouse - books in high shelves, on pallets, everywhere. Often hundreds of issues of the same one.

      But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.

      Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.

      We were taught respect for books, but not everything is worth preserving.

      • ghaff6 hours ago
        I didn't go this year but my local town library has a book sale where books go for something like $10/bag. I donate some books to them throughout the year. There are still a lot of books available on the last or second to last day of the sale. I'm sure a huge number get pulped.
    • sly0106 hours ago
      Well, they could turn the bad faith story into a good faith story by making them available for everyone to download perhaps. (AI companies "saving" old books!) But that would require giving a s*t which they don't and that is the real problem imho.
      • Aurornis6 hours ago
        They legally cannot do this.
        • sly0106 hours ago
          That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

          People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.

          • Aurornis5 hours ago
            > That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

            I don't think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That's equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what's needed to make this happen would be five figures per book to get started.

        • theroadnotbacon6 hours ago
          That certainly hasn’t stopped them before… IP theft is kind of their whole thing, isn’t it?
          • Aurornis6 hours ago
            Distributing copyrighted works (prior to expiration of their copyright) verbatim is illegal.

            Training an LLM on copyrighted works is not illegal.

            This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.

          • syrrim5 hours ago
            Google attempted to do this 15 years ago, they got sued and stopped. It turns out that tech companies occasionally do have to follow the law, you'd think people would be happier about that...
            • keeda5 hours ago
              Wait, if you’re talking about the Google Books case, Google won. Maybe they made adjustments on how they served results but they certainly did not stop.
              • ndiddy3 hours ago
                The Google Books settlement was originally going to make Google into a clearinghouse for scans of out-of-print books. The scans would have been available for individuals to purchase for a reasonable price, and libraries and institutions would have been able to subscribe to a service that would give patrons access to the full text of every book. This deal fell apart because some research libraries and authors argued this was anti-competitive, as anyone wanting to make a competing service would have to go through the same process as Google of settling a class action lawsuit. They instead wanted Congress to pass a law to free up the rights to orphaned books. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they're out of print when ebooks and print-on-demand exist is that they won't get enough sales to make it worth the time and money to figure out who the royalties should go to. The result is that nobody outside Google gets to see the full Google Books scans.
        • 2OEH8eoCRo05 hours ago
          They legally cannot scan them in entirety either but they are.
          • ChickeNES5 hours ago
            Again, under Bartz v Anthropic they can scan and train on whatever they want, as long as the original is lost in the process.
    • laybak6 hours ago
      I have a similar thought too each time I walk past piles of discount books.

      I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are

    • a_shovel5 hours ago
      A lot of these books are reference titles that are outdated to the point of uselessness and/or weren't that great/interesting when they were new. Not many people have interest in or use for textbooks from the 50s.
    • icantevenhold6 hours ago
      Why do they shred them at all instead of donating or selling them again?
      • andrew_lettuce6 hours ago
        My understanding is they cut off the binding for scanning. They'd have to resell it donate by the page
        • ChickeNES5 hours ago
          They have to destroy the copy either way for it to be fair use (according to Bartz v Anthropic). Cutting the bindings off is already the faster method to scan, and once you're required to pulp the original anyway, it becomes a no-brainer.
          • shagie5 hours ago
            The wording doesn't exactly say that. https://www.akingump.com/a/web/h6WFidTYyTnoNEPehPXMYu/auvY7D...

            The relevant part of the ruling starts on page 27.

                For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
                The third fair use factor favors fair use for the purchased library copies converted from print to digital.
            
            Fair use favors the destruction to avoid accidental surplus copying of the source material. It doesn't require it. If Anthropic put everything in a warehouse, they could keep it in the warehouse... but they couldn't do anything with them afterwards. They couldn't sell them as books or donate them as that would mean that it wasn't fair use and copyright infringement would have taken place once that additional copy was distributed again. At that point, fair use and economics both favor destroying the original.

            If I make a DVD copy of an old VHS tape that I own, that's fair use format shifting. I can keep the VHS tape without issue. I cannot donate it to the library or put it out in a garage sale. When I got rid of my VHS player, I threw out VHS tapes too since they were of no use and only cluttered my shelf.

      • quickthrowman6 hours ago
        A judge ruled it was OK to scan books and save the scanned copy if you shred the physical book afterwards.
        • shagie6 hours ago
          The judge ruled that format shifting was fair use. The fair use argument was enhanced because the original was destroyed. Hypothetically, one could do only the format shifting and keep the original - but the original could not be sold or donated or otherwise given away because then the format shifted copy would not be fair use.

          I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.

    • pixl976 hours ago
      Yea, there are a ton of people that seem they'd rather the books get lost forever than be looked at by an AI company.
    • palmotea6 hours ago
      > Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

      Yeah, now instead of old books being turned into toilet paper, we'll get turned into toilet paper. Much better.

      But Sam Altman will become richer than God, and isn't that what really matters?

      But don't worry! You'll still have access to ChatGPT until your savings run out.

      • ChickeNES6 hours ago
        Do you have an actual point about scanning the books, or are you just using this as a soapbox to rant about AI and Sam Altman?
        • palmotea5 hours ago
          Yes, you missed it. Perhaps you should read it again until you get it?
          • ChickeNES5 hours ago
            Again, do you have an actual objection to the subject of the article, or are you just mad it enriches people you don't like?
            • palmotea4 hours ago
              > Again, do you have an actual objection to the subject of the article

              I was responding to a comment, which apparently is another thing you missed. Maybe work on not doing that, instead of snarkily asking for everything to be carefully spelled out to you?

              > or are you just mad it enriches people you don't like?

              Nice strawman, pity if someone knocked it down.

              • ChickeNES4 hours ago
                Okay, so all you have are personal attacks, goodbye.
                • palmotea2 hours ago
                  > Okay, so all you have are personal attacks, goodbye.

                  Note: I made no personal attacks.

  • patall6 hours ago
    Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?
    • Legend24406 hours ago
      It sounds like it is about the information in the books. The titles they're looking for are all nonfiction. Not everything is on the internet, and just because it's a few years old doesn't mean it's outdated.

      Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.

      The goal here is to have all human knowledge in a single file, which is pretty neat IMO.

      • newsy-combi6 hours ago
        The internet basically never delivered on the promise of replacing textbooks or even education as a whole. Wikipedia sucks on many topics, has insane internal politics, and is a tertiary source by design (redigesting blogs and books), whereas textbooks are generally secondary.
      • patall2 hours ago
        I see. So it is much less about 150 year old fiction books, but more about 1980s science literature that was only ever printed five times. That makes a lot more sense than what the public debate seems to be about.
      • ChickeNES6 hours ago
        > the information density of published books is a lot higher than most internet text

        I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.

        • criemen6 hours ago
          They're targeting non-fiction books, so romance novels would be out.
          • ChickeNES5 hours ago
            Booksellers noticed a huge uptick in non-fiction purchases, that is not the same thing as them not targeting fiction at all.

            Edit: Actually, I have real evidence, the Bartz in "Bartz v Anthropic" is Andrea Bartz, a novelist, and the complaint specifically lists four of her novels as infringed works.

          • scottyah6 hours ago
            Is that a policy change after o4 got a little out of hand?
      • zardo6 hours ago
        Also the ability to set the training input limit in the past could be useful.
    • hyperhello6 hours ago
      They are the memories of the productive part of society. You can leaf through them and get the feel of what it was like. You don’t need most of your memories, personality, or core ideals to be productive to the State.
    • Aurornis6 hours ago
      > but now a few thousand books are what is needed to run a successful AI company?

      They're scanning millions of books.

      It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.

    • layer86 hours ago
      Diversity. Modern books with modern content and in modern styles are overrepresented, old ones underrepresented.
    • acuozzo3 hours ago
      > Can someone explain why the old books could really be relevant.

      Lots and lots of information is not online. You'd be surprised.

    • timcobb6 hours ago
      I'm guessing the more integration tables they consume, the better they become at integration.

      I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.

    • ijk6 hours ago
      Because Yahoo in particular was very good at destroying goldmines shortly before they became ultra valuable.

      There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.

    • theroadnotbacon6 hours ago
      I also wonder if it’s used for text generation in image models! Awful lot of typefaces, sizes, orientations, and words in those books.
    • cesarvarela5 hours ago
      I think a book is like one completion of the mind behind it, so in a way, this is just distillation.
    • aprilthird20216 hours ago
      The information density is a lot less for millions of yahoo groups. They are far more likely to cover the same topics and not have new information in them.

      Books are more likely to be about a specific topic or story or time or setting and be more information dense

    • nemomarx6 hours ago
      Writing style maybe?
    • 2OEH8eoCRo05 hours ago
      They aren't but these companies have more money than sense.
    • mistercheph6 hours ago
      > why the old books could really be relevant > it can barely be about the information in those books (that would be very often outdated),

      LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.

      They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).

    • aisenik6 hours ago
      Cognition is encoded in language, they weren't brain-damaged yet. There's better (real, not token-exchange) thinking, which LLMs can copy and reproduce in novel arrangements.
    • mars-or_wars6 hours ago
      [dead]
  • RobotToaster7 hours ago
    The sad part is, imagine how positive this could be if the scans were made available to the public.
    • IrishTechie6 hours ago
      Might add insult to injury for the publishers/authors though?
      • RobotToaster6 hours ago
        Maybe for current in print books, but the concern is about rare books being pulped in this process. If a book is rare then it isn't in print, so nobody is making money out of it.
        • nightpool6 hours ago
          Unfortunately copyright law does not have a squatters-rights exception
          • ChickeNES6 hours ago
            I've believed for years that it really should have one, or at least a "you aren't selling this to the public at a reasonable market-rate price (or via subscription, I care about access far FAR more than ownership), you lose all rights to it" regime.
      • doublerabbit6 hours ago
        A book, song that has been out in public-domain for more than 10 years, should be downloadable. Even if you were to rebuy the same CD again the amount that the artist received would be a pointless pittance. Why not just let it be free to be enjoyed by all?

        OCR Scanned for training, then tossed away or burnt. Great for nature.

        • ChickeNES5 hours ago
          Why would they be tossed or burnt? Tons of discarded books are simply recycled like any other paper object (that's why people keep saying they were "pulped")
  • rjh296 hours ago
    Good memories of visiting used bookshops near Kyoto University with stacks upon stacks of obscure literary works and research material. Lots of interesting books about the Japanese language that were never digitized and I always left with 2-3 new books. So I'm not super happy about AI companies hoovering this all up and not making the scans available.
  • YVoyiatzis6 hours ago
    Pretty much as it happened with vinyl records twenty years ago. I remember seeing photos of this guy somewhere in Brazil standing atop heaps of vinyl records which he had amassed with HDLR intention. Now books. Some of us hold on forever.
    • mistercheph6 hours ago
      [flagged]
      • ChickeNES6 hours ago
        "the engine of human progress"? I think you mean capitalism.

        > the books being burned by these misanthropic lunatics are not available in any other medium

        prove it, name one title

        > this is not about fascination with some particular mediumn of transmission

        it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.

  • shooa day ago
    Via google translate:

    > It has been discovered that online used bookstores across Japan have been receiving a surge of large orders for books since around August of this year. Interviews with these bookstores reveal reports of "100 books sold per day" and "days where sales have increased fivefold," leading to widespread speculation within the industry that the orders are intended to collect training data for generative AI (artificial intelligence). Further investigation revealed records of over 50 tons of books being exported from Japan to the United States. Is it acceptable for books to be consumed and discarded for AI training?

    [...]

    > Nippon Television investigated using "Sayari," a tool that analyzes import and export data, and confirmed records that a group company of this major Japanese book distributor exported more than 50 tons of "JAPANESE BOOKS" to the US since last year. Assuming that all the books were heavy hardcovers (calculated at 500 grams), this would amount to the equivalent of 100,000 books.

  • dofm4 hours ago
    AI firms buying and destroying the sources of knowledge is a weird way to get us to Fahrenheit 451 but maybe the USA will get there before it switches to metric after all.
  • karim79a day ago
    Buy as many books as you want but please don't burn after scanning.
    • _ink_a day ago
      Unless somebody steps in and prohibits it they will do it, because somehow that's most cost efficient.
      • karim79a day ago
        Yeah. It somehow keeps the competitors at bay, apparently, from previous discussions about AI companies who burn books.

        Can we really trust these AI companies when they're basically assimilating human art and culture and everything, only to go on and destroy the evidence?

      • I don't understand why can't they donate the books somewhere. Even if they don't want to share the scans, why not give up the books?
        • piva00a day ago
          In the name of efficiency they cut the spines, rip the pages apart for easier scanning, books are literal dead wood which would cost a lot in shipping so it's easier and cheaper to just discard or burn the remnants.

          The book ceases to be a book right at the beginning of the extractive process.

        • secabeena day ago
          Largely, because books are heavy, and there's limited demand for most used books. If you've ever been at a book sale on the last day, there are always lots of leftover ones that no one wants.

          There's a general sense that librarians are preservationists. This is far from the truth. Most books get pulped within a few years of printing. Librarians are constantly discarding books that don't get checked out to make space for new books.

        • EA-3167a day ago
          They can, but it would cost more and take longer, so they don't.

          It's not as though they particularly care, and it's not as though bad press matters with a captured regulatory environment.

    • itmma day ago
      Isn't the "burn after scanning" due to copyright? It seems like destructively scanning was an explicitly allowed method in Bartz v Anthropic as transformative fair use.
      • 21 hours ago
        undefined
    • What do you want them to do with the books after they’ve digitized them?
      • karim79a day ago
        Give them away. Put them on the street for people to pick up. Donate them to anyone. I can't tell if you're being serious or if your comment is sarcasm. I'm leaning toward the latter.
        • You want them to put them on the street in front of the warehouse? Do they post them on free stuff Facebook group after or just wait for the garbage truck?
  • glimshea day ago
    I'm strongly against destroying books to train AI. That said, I'd be extremely interested in exploring the contents of Japanese books, which probably have a lot of stuff not available outside Japan, through a LLM.
  • t1234s7 hours ago
    The value in these AI companies will be more in their proprietary training data than the models.
  • ivanjermakova day ago
    So they ran out of information on the internet...
    • They ran out of free information on the internet. They can't steal it just like that anymore because the providers are now aware of the value they could provide and could sue the now trillion dollars companies for stealing the data.

      They didn't bother when it was a non-profit sponsored by Elon Musk.

    • r_leea day ago
      or more like they can't use the internet anymore because it's just slop now
  • mistercheph5 hours ago
    Before you buy the apologia that these books are not valuable or interesting: if they weren't valuable or interesting the AI labs would not be spending billions of dollars to purchase and scan them. Yes, everyone has a personal anecdote about pallets of garbage books but I have three insights for you that you may not have because you don't read or sift through pallets of garbage books:

    1) The labs don't want garbage books, they want interesting books that are rare and unique. They want high quality training data, random permutations of language style are fine, but what you want is unseen information, unseen patterns of thinking, unseen ideas.

    2) Most pallet of books contains lots of valuable and interesting works, maybe 1-3% but sorting through them takes time, money, and energy, that's why the labs are starting to purchase by the pallet, it's because they already have a fully automated process so they can always beat any bookseller small or large on cost to find the books of interest and value in a pile.

    3) Many of these pallets may sit for years before being sorted, and many of the books may sit for years before being sold, but these things actually do eventually happen, valuable books are found, and they eventually make their way to interested readers, this is the business model of used bookstores. Most books of value don't get destroyed or thrown away.

    Destroying human art, knowledge, and culture is an essential part of the business plan for frontier labs, it is not enough to steal and regurgitate all the art and information in the world, you also want to make it inaccessible through any other means than the regurgitation machine. Don't expect the book burning to be an isolated incident, they are coming for every other form of stored human knowledge or art, and yes, unfortunately while scanning it they will have to destroy the original copy. And attacking the past is only the beginning.

    • ChickeNES5 hours ago
      I have sifted through everything from piled up junky independent book shops, to cast offs from research libraries, to dumpsters of end of life books. Most books really are not worth the paper they are printed on.

      > Most books of value don't get destroyed or thrown away.

      Of value to who? Most used bookstores are boutiques that over-curate and will happily refuse or recycle books that they deem are inferior/irrelevant. This is a big reason why I prefer Half Price Books over most any other used bookstore, they sell most everything.

  • panny7 hours ago
    A dark age will come. AI shredders destroy all the books then hallucinate what they once contained.
    • manarth6 hours ago
      Once upon a time, Hansel and Gretel were walking through the woods when they met a Sleeping Beauty called Snow White. As they tried to wake Beauty, a naked Emperor walked in screaming "Off with his head" before a Big Bad Wolf started huffing and puffing.
    • aprilthird20216 hours ago
      I'm the opposite of an AI doomer but this is actually scary to me. Once knowledge is hollowed out like this how can we get it back?
      • MarkusQ6 hours ago
        We could go out in the world, have experiences, cogitate upon them, learn to write well, and then do so? It ain't easy, but that's the way it used to be done.
      • mistercheph6 hours ago
        This is unironically part of their business plan: don't just regurgitate the world's information but also destroy all other sources of information.
  • Haven88014 hours ago
    2000 years ago Qin burnt books. Barbaric. Then Hitler burnt books. Barbaric. Now AI burnt books. Is legal and most cost effective ways. When Skynet happened, hard to pity barbaric humans.
  • 6 hours ago
    undefined
  • a day ago
    undefined
  • Armecera17 hours ago
    [dead]
  • fangspire6 hours ago
    [dead]
  • swingandamiss7 hours ago
    [flagged]
  • tsylba6 hours ago
    Ah yes, the litteral destruction of culture and physical media for a centralised subscription service. I love the liberal world of techno enclosures of our new overlords, viva el free market economy.
    • quickthrowman6 hours ago
      Feel free to buy books by the ton and preserve them yourself. The simple fact they’re being sold by weight implies they’re not rare or unique.
      • mistercheph6 hours ago
        The simple fact that the AI labs are spending billions of dollars to acquire and scan them implies they are rare and unique.
        • quickthrowman5 hours ago
          Please provide proof for the billions of dollars claim, thank you.
    • Analemma_6 hours ago
      What do you think happened to all these used books before the AI companies showed up?
      • alightsoul6 hours ago
        They sat in boxes
      • ghc6 hours ago
        "Once great literature—now great litter."
  • duchanjo7 hours ago
    Could this be for AI training data?
    • ErneX7 hours ago
      It’s on the headline if you visit the link.
  • genxy6 hours ago
    It doesn't matter what language the tokens are in, now Eye of Sauron seeks to consume all knowledge.
    • pfdietz6 hours ago
      How does one "consume knowledge"?
      • wccrawford6 hours ago
        Well, in this case, I imagine they mean by shredding books after scanning them. Since that's what's happening.
        • pfdietz6 hours ago
          That's not consuming knowledge, that's consuming cellulose and ink. If it were, then printing another copy of a book would be "producing knowledge".

          BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.

          • genxy3 hours ago
            When the refcount goes to zero the knowledge is consumed. Your take is overlay pedantic without benefit.

            Both definitions of knowledge creation can used. If one creates a book and it is never read, has it been produced? Isn't knowledge also its access and how widely it is disseminated?