

If you own a book, you can legally make as many copies of it as you want. As long as you’re not distributing them, the courts treat that as a single effective copy.
Father, Hacker (Information Security Professional), Open Source Software Developer, Inventor, and 3D printing enthusiast


If you own a book, you can legally make as many copies of it as you want. As long as you’re not distributing them, the courts treat that as a single effective copy.


Legal reasons: When you “purchase” an ebook you’re actually just licensing it and nearly all ebook licenses exclude the ability to do anything with the ebook other than read it yourself.
They could get a commercial license to get big ebook libraries like Anthropic did, but they can’t, really, because Anthropic paid for an exclusive license. Which means that if other AI companies want to compete, they sort of have to buy books in bulk and tear them apart to scan them.


If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.
Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.
It’s ok to throw trash away! Really!


That’s like saying, “they had some failure modes from the synthetic data, so they should just obviously stop trying forever.”
They’ll just fix the edge cases and move on. Like any programming task.


Yes. That does make sense.
If you bought a book, scanned it—destroying it in the process—then read it on your computer, that would be completely acceptable.
Why is it wrong when a corporation does the same thing?
They’re not claiming ownership of the copyrights, just ownership of a copy. Which is how copyright works.


Big AI mostly switched to synthetic training data anyway. The books they’re digitizing are being used to gather knowledge, not writing styles or logic (mostly).
As in, when you ask ChatGPT how long some book is, it can just go check (if it’s in the database). It’s also useful if you ask about that book or about knowledge contained in that book. It’ll even reference books now (if you demand that in your prompt).
It’s not the same as earlier LLM tech which relied on scanned text to figure out how to respond to any given prompt (from a language standpoint). The “language” part of LLMs is a solved problem now (thanks to the synthetic training). At least for English 🤷


Is it really destroying though? They’re digitizing them, and publishers still have the digital copies ready to print more at any time. So it’s not like they’re destroying the texts, they’re just shifting them.
Nobody complained when Google did this over a decade ago 🤷
When you say they’re “destroying the books” you make it sound like they’re erasing one of the last known copy of some important work when in reality, most of these books were purchased in bulk from bookstores and libraries that were planning on discarding them anyway.
Almost all these books were either headed to the dump or the recycling center. They’re just being digitized on the way.
If you want unautomatable, ludicrously expensive, insecure mountains of instability… Go for it👍
Proving once again that windows isn’t ready for the desktop.
The chances this fixes it are… Remote.


While I sort of agree with you that that is his angle, service jobs that aren’t easily automatable will continue to climb in the “job stability and pay rate” category until the world figures out how to deal with Baumol’s Cost Disease:


The AI paradox: It’s both original (hallucinating) and plagiarizing (copying things, word-for-word).


This is why open source AI is necessary!


There is a story people tell about AI regulation, and it goes like this: the technology is moving too fast, governments can’t keep up, regulators are overwhelmed, and by the time anyone writes a law the thing they’re trying to regulate has already evolved into something else entirely.
No. That’s not the story people are telling about AI regulation. It goes like this:
If we regulate AI, that will give an advantage to AI companies in other countries. They will surpass our AI capabilities and leave us in the technological dust.
There’s a related story:
If we regulate AI, we’re likely to create more problems because Boomers don’t understand technology.


Everyone wants to access Netflix, YouTube, Prime Video, etc through their TV interface and I just don’t get it. The best experience is when you hook up a PC to your TV… not some TV-centric Android OS or Roku’s thing.
Install Kubuntu on some old PC with a GPU that can handle 4K @60Hz and you’re good to go. KDE and Firefox let you crank up the zoom so everything’s easy to read and it even has HDR support (though I prefer going without it… Old person eyes).
It’s such a vastly superior experience. Not only do you get the usual stuff, you can use a real keyboard to type into that search bar. You can also access all those pirate streaming sites and do normal PC stuff like play games.


I browse “All” most of the time and the Femcel Memes show up pretty regularly. It’s not like this gif though. It’s more like watching a super interesting science experiment.
“Ah, yes. I see. I see. How interesting!”
My scientific notes so far:


Zuck is having all his keystrokes recorded too, right? Right?
Make sure the AI engineers get that data. Especially the passwords to his and the company’s bank accounts. All his accounts, actually.
There’s a reason why most businesses don’t implement keystroke logging.


The reason why you hear this so often is because academia is designed to teach students based on a logical, reasonable curriculum. The curriculum will be mostly well-thought-out and cover all the important topics.
Then you take someone who followed this perfectly reasonable path and you place them in front of the total shitshow that is most businesses. Everything they think they know won’t be applicable because most of the time, logic and reason were not what drove adoption of any given tool or practice.


I’ve been researching this a bit… I’ve come to the conclusion that there is no AI bubble. In fact, we’re only just getting started down this road. Unless there’s some massive 100x efficiency breakthrough in training AI and inference, the entire world is going to be building seemingly endless AI data centers (and the normal compute kind, e.g. for stuff like AWS, Google/YouTube, Meta, banks) for at least a decade. Probably a little longer (12-15 years before demand levels out).
Everyone thinks that “AI data center” means ChatGPT, Claude, Gemini, etc but there’s 10,000x more demand for AI than those services. Think: Pharmaceutical companies trying to find proteins, scientists (and big agriculture!) trying to model the weather, and other businesses trying to automate stuff. Not just software; robots and things like conveyor belts.
Another example: Ever use one of those self-checkouts that’s mostly just a camera pointing down, where you place the stuff you’re purchasing? That uses AI too.
Having said that, there is a great big bubble in AI: OpenAI, specifically. That will definitely pop one day. And hopefully, the DRAM bullshit will go along with it.
Haha: What you state is logical and reasonable. That position would be easy to defend, if the legal system around copyright made sense.
Unfortunately, it isn’t a logical system. You’re trying to apply copyright law—which has now been ruled on in court, officially making training AI with copyrighted material Fair Use.
The problem is that the issue with ebooks is all about contract law. Not copyright.
When you “purchase” an ebook (which is a misnomer), you’re actually signing a legally binding contract. A license, in effect, to use that copyrighted work for the sole purpose outlined in the contract. That outlined purpose expressly forbids using the work for anything other than you—the purchaser—reading it (usually on the platform they specify).
Having said that, many, many court rulings have found numerous license clauses like that to be unenforceable. That is, just because it’s written in the contract, doesn’t mean it’s legal.
The courts have ruled—thanks to everyone’s efforts fighting the MPAA, RIAA, Sony, Microsoft, and Nintendo in the 1990s and early 2000s—that everyone does have the legal right to “platform shift” whatever copyrighted works they own.
To get around that, those very same entities tried to use the Digital Millennium Copyright Act’s rules about circumventing “copyright protection mechanisms” to try to make it illegal for people to platform shift their stuff anyway. That is, they added trivial encryption to all their platforms.
But it’s even more complicated than that! Because Congress gave the Librarian of Congress the power to say when it’s legal to circumvent such “copyright protections.” For example, technologies that aid the blind (I.e. gotta decrypt that file for the program to read it out loud).
There’s other exceptions and everyone fights to get more added whenever it comes up.
The key takeaway, though is that none of this complicated mess applies to physical books! So there ya go 😁👍