Need help?
<- Back

Comments (276)

  • thread_id
    I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.https://en.wikipedia.org/wiki/Google_Bookshttps://arstechnica.com/tech-policy/2015/10/appeals-court-ru...https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
  • throwaw12
    America is interesting.* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz* download and publish a book as a company -> fine 1% of revenue* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model
  • ziyadb
    I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
  • cladopa
    It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
  • bix6
    > A guest post by Anna’s Archive volunteer “u” (translated from Chinese).I do not know much about Anna’s. Is it Chinese or is this just one of many worldwide helpers?
  • SquireBuilds
    Do you think they will make all of that publicly available after it's been scanned?
  • akk0
    I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
  • pmoriarty
    Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
  • CamelCaseName
    You ask "Why destroy physical books?"I ask "Why save physical books?"If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
  • godber
    I think a solution to this could be for the government to require the companies to provide the original scans and OCR results to the government for safe keeping until copyright expires.Clearly that makes assumptions about the function of government and the complacency of copyright holders.
  • twright
    I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.
  • branon
    I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the listMuch like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
  • clarionbell
    In my experience, when someone mentions burning library of Alexandria, they are either exaggerating, or have a poor grasp of history. Usually it's both. This post, the discussion, do not change my mind.
  • shrubble
    AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
  • ryandvm
    It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
  • thisisauserid
    Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
  • lukasbm
    A more honest framing would be: "The judge ordered the destruction during scanning because of stupid copyright laws"
  • ZoomZoomZoom
    The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
  • juiceland
    Why are people acting like books can’t be reprinted?
  • michael0church
    What’s really disgusting is how unnecessary this is.LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.
  • thuuuomas
    Is there any evidence the books are truly "destroyed" & not merely "disassembled"?It's common practice to cut the spine & binding off a book, scan the loose pages, & drop the rubber-banded loose pages in a box somewhere. The book still exists, just without its binding.
  • lo_fye
    The price they should have to pay for destroying a rare book (just to scan it) is making a pristine high resolution digital copy of it available to the public at no cost whatsoever.
  • 1970-01-01
    Again, rare and valuable are not the same thing. A $2 bill is not as valuable as you think it is, unless you think it is $2.
  • infecto
    I genuinely have yet to connect on this idea that we are “burning Alexandria” or AI companies are ruining the future of humanity because most of these books are absolutely junk.I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.
  • JsonDemWitOster
    Sorry for the copy-paste from elsewhere. I'm probably doxxing my online accounts with this. 100% my words tho and 100% I stand by this.Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
  • bhouston
    Why not force them to release their records? Some legislation would help. If they are scanning the world's books, the results should be open.
  • luciana1u
    the irony of scanning rare books to preserve them while the scanning pipeline is what's destroying the physical copies is going to be a great trivia answer in 50 years
  • bethekidyouwant
    I don’t get this latest anti AI talking point. They are digitizing the books preserving them forever.. are you upset that you don’t have access to it? Because you didn’t before either… stop whining and give AA some money.
  • ionwake
    I mean guys i get it, you have a bad day just once and you want to rewrite history, we all do, but geez... I guess just I feel destroying books should be classed as unethical. if i was ruler I think id ban it.
  • josefritzishere
    Destroying history is a crime against humanity.
  • r721
  • carlosjobim
    I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
  • ForHackernews
    Project Unica is an initiative by the University of Illinois Libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...
  • RcouF1uZ4gsC
    What often gets missed is that they are buy one physical copy and turning it into a digital copy.They have done zero to destroy the durability. In fact, it’s probably more durable.If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
  • Razengan
    Ideally, governments or international organizations should be doing this: "harvesting" all the media output by humanity and making it available for everyone, similar to the Library of Congress etc.Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.
  • starkd
    How is Anna's Archive getting around the copyright violations of hosting all these books for access to all? I suspect it won't be long before they get sued and are forced to shut down. I spent some time reading the web site, and it doesn't look to be a well thought out project. Even the way it is organized leaved much to be desired. There's much more to library science and the organization of a vast collection of books than meets the eye.
  • next_xibalba
    Why are these books "rare"? Because no one wants them. Why then are we up in arms over their destruction? These sensational headlines make it seem as though a copy of the Codex Sassoon 1053 is being destroyed, when in fact these are just obscure books that no one cares about.The rhetoric on this topic is reminiscent of the rhetoric regarding data centers: some noxious combination of misinformation, misunderstanding, and sensationalism, wielded against technological progress.
  • shevy-java
    But scanning the books also helps those AI companies because ultimately they want more data. Yes, they also destroy rare books to sabotage competitors, and thus also damage global society - a reason why these evil companies should be disbanded - but the article seems to not put any thoughts into things here, other than the superficial "they destroy books".
  • greenavocado
    During World War II and its immediate aftermath, between 35 million and 40 million books were destroyed in Germany due to Allied actions
  • m00dy
    Since when books have become a supply limited asset ?
  • eulgro
    We've been seeing that headline for a few weeks now and I really don't understand the problem.Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.So what's the problem here exactly?Also from the article:> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
  • warkdarrior
    > Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
  • maxlin
    It is quite disappointing to see them not using the type of machines that don't actually destroy the books, like, afaik, Internet Archive is using.Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
  • spwa4
    The problem is the choice made here: this is the world's 2 major governments choosing to give very large legal advantages to AI models, over actual people, in copyright. US and EU governments obviously want AI models to make everything from books to movies in the future, and this is a conscious choice both governments are making without consulting people.Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:a) EU companies making ML models have to self-sabotage against their competition.b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.[1] https://commission.europa.eu/document/download/f0b8d4c3-51aa...[2] in their source you can see what models they were likely using internally 2 years ago: https://github.com/openeuropa/gpt-at-ec-php-client
  • sudeepsd__
    [dead]
  • anon
    undefined
  • grammarisking
    [dead]
  • maxdo
    Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court