Short-term boom could lead to a longer-term drought and removal of important works from circulation.
130 comments
I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.
But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.
Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.
We were taught respect for books, but not everything is worth preserving.
Most of the information in it is practically worthless today unfortunately - Saddam Hussein still alive, as were Katharine Hepburn and King Hussein of Jordan, no mention of Ceres and Pluto was still a planet, and the Internet and software had barely a mention. Most articles had less information than a standard Wikipedia article.
That being said, the way information was presented in them still outshines anything one may find on the internet today. Even the simple elements - neat diagrams and relevant images, proper sectioning and organization of the text, footnotes to other relevant articles...
Long gone are the days when I would simply take a bowl of ice cream and a volume and just read it end to end.
I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are
Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.
The goal here is to have all human knowledge in a single file, which is pretty neat IMO.
I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.
They're scanning millions of books.
It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.
LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.
They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).
OCR Scanned for training, then tossed away or burnt. Great for nature.
> It has been discovered that online used bookstores across Japan have been receiving a surge of large orders for books since around August of this year. Interviews with these bookstores reveal reports of "100 books sold per day" and "days where sales have increased fivefold," leading to widespread speculation within the industry that the orders are intended to collect training data for generative AI (artificial intelligence). Further investigation revealed records of over 50 tons of books being exported from Japan to the United States. Is it acceptable for books to be consumed and discarded for AI training?
[...]
> Nippon Television investigated using "Sayari," a tool that analyzes import and export data, and confirmed records that a group company of this major Japanese book distributor exported more than 50 tons of "JAPANESE BOOKS" to the US since last year. Assuming that all the books were heavy hardcovers (calculated at 500 grams), this would amount to the equivalent of 100,000 books.
Read the full thread on Hacker News →
Related stories
- Hacker News · 3 points · 8 days ago
- Show HN: A bookshelf of rare company documents and out-of-print tech booksrare-books.vercel.appHacker News · 10 points · 3 days ago
- Hacker News · 2 points · 7 days ago
- Hacker News · 1 points · 6 days ago
- QuestDB (YC S20) Is Hiring a Sales Engineerquestdb.comHacker News · 1 points · 8 days ago
- The Library After Dark – a walkable 3D library of 600 public-domain bookslibraryafterdark.spaceHacker News · 2 points · 5 days ago