I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.
But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.
Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.
We were taught respect for books, but not everything is worth preserving.
Yeah, now instead of old books being turned into toilet paper, we'll get turned into toilet paper.
And Sam Altman will become richer than God, and isn't that what really matters?
But don't worry! You'll still have access to ChatGPT until your savings run out.
I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are
OCR Scanned for training, then tossed away or burnt. Great for nature.
> the books being burned by these misanthropic lunatics are not available in any other medium
prove it, name one title
> this is not about fascination with some particular mediumn of transmission
it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.
Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.
The goal here is to have all human knowledge in a single file, which is pretty neat IMO.
I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.
They're scanning millions of books.
It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.
There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.
I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.
Books are more likely to be about a specific topic or story or time or setting and be more information dense
LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.
They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).
> But what happens when sales numbers don't meet projections? The book is discounted. Then, at the publisher's discretion, the bookstore will receive a directive to rip the covers off the books, recycle the remainder of the book to be "pulped" or turned into other forms of paper, such as notebook paper and toilet paper. The bookstore is expected to mail the book covers to the publisher as evidence that the book has been destroyed.
BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.
I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.
People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.
I don't think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That's equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what's needed to make this happen would be five figures per book to get started.
Training an LLM on copyrighted works is not illegal.
This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.