

If it’s part of a fingerprinting lib, then they just fucked it up. Because the fingerprinting can be done in one ~16ms window of time when the script runs and then you can break down the audio objects and pretend nothing happened.
Father, Hacker (Information Security Professional), Open Source Software Developer, Inventor, and 3D printing enthusiast


If it’s part of a fingerprinting lib, then they just fucked it up. Because the fingerprinting can be done in one ~16ms window of time when the script runs and then you can break down the audio objects and pretend nothing happened.


Honestly, this doesn’t sound like spyware. More like absurd optimization. Here’s why: Audio devices often go to sleep (to save battery) and can take a second or two to wake up. That’s enough time that the end user will miss the first second or two of audio.
By keeping the audio outputting something (even if it’s just silence), they can guarantee that their little (often hilariously terrible) product videos will play the way they expect (which is loud and startling, of course!).
This is just one of those stupid tricks that’s bad for energy use but good for ignorant users who might complain that the first second of every AliExpress video is silent 🤷
The reason why I believe this to be the case is because there’s really nothing unique or interesting to be learned from an audio output loop that’s literally sending zeros through itself. If you wanted to use that to fingerprint a user, you wouldn’t need to keep it active. You could pass a single zero (silence) through and be done.


Shit: Is this why people ask me random IT questions and suddenly vent about their lives in chat?


LOL! I was making a joke.


At least they’re being honest about it. 99% of the companies training AI or selling people’s data to AI companies are doing so without telling anyone.
Having said that, Twitch streams are public almost all of the time so complaining that they’re being used in ways the creator never intended is a bit like complaining that someone caught you picking your nose in public.
We learned this lesson in the 90s: Once something goes public on the Internet, it is no longer under the creator’s/publisher’s control. People can and will do whatever TF they want with it. As long as they’re not redistributing it without permission, that is legal.
It’s everyone’s right to do whatever TF they want with the stuff that’s literally sent to their computer on purpose. That includes using that thing to train AI.


Yet another comment that says they solved the problem without saying how.


It’s beyond “an arms race.” AI has become a commodity that’s everywhere and it’s being used and developed by zillions of completely different groups all over the world with wildly varying use cases and completely different goals.
At this point, trying to ban it would be like trying to ban dandelions from growing.


For shore, you pelican’t fish with them around.


Someone cut above the rest‽


Massive overkill!
That’s a great Home Theater PC. Hook it up to your TV and install KDE Plasma (distro of your choice 👍) then just crank up the global desktop zoom to like 225% and get:
Now install Docker and all the containers/servers you want!
I have a ~14yo HTPC with a mostly modern GPU (8GB) that only has 8GB of RAM and four cores on some old AMD CPU. It’s running four separate servers for my self-host purposes and it doesn’t break a sweat.


I did the math and these are the approximate real numbers:
That’s active Linux desktops. That is, they’re being used to browse the Internet regularly enough to make it into StatCounter.
They’re just like any other data center, except bigger because that’s more efficient than building a bunch of smaller ones.
AI data centers have a lot of GPUs and specialty hardware that use a lot of power, but they’re still just big buildings with lots of computers in them. They just sort of sit there, with air conditioners running 24/7.
If they’re not the huge “hyper scale” type, you end up with a situation like in the article: Where they had to cram too many cooling units in too tight a tight space (to make up for the fact that the data center wasn’t built for its current workload).


Haha: What you state is logical and reasonable. That position would be easy to defend, if the legal system around copyright made sense.
Unfortunately, it isn’t a logical system. You’re trying to apply copyright law—which has now been ruled on in court, officially making training AI with copyrighted material Fair Use.
The problem is that the issue with ebooks is all about contract law. Not copyright.
When you “purchase” an ebook (which is a misnomer), you’re actually signing a legally binding contract. A license, in effect, to use that copyrighted work for the sole purpose outlined in the contract. That outlined purpose expressly forbids using the work for anything other than you—the purchaser—reading it (usually on the platform they specify).
Having said that, many, many court rulings have found numerous license clauses like that to be unenforceable. That is, just because it’s written in the contract, doesn’t mean it’s legal.
The courts have ruled—thanks to everyone’s efforts fighting the MPAA, RIAA, Sony, Microsoft, and Nintendo in the 1990s and early 2000s—that everyone does have the legal right to “platform shift” whatever copyrighted works they own.
To get around that, those very same entities tried to use the Digital Millennium Copyright Act’s rules about circumventing “copyright protection mechanisms” to try to make it illegal for people to platform shift their stuff anyway. That is, they added trivial encryption to all their platforms.
But it’s even more complicated than that! Because Congress gave the Librarian of Congress the power to say when it’s legal to circumvent such “copyright protections.” For example, technologies that aid the blind (I.e. gotta decrypt that file for the program to read it out loud).
There’s other exceptions and everyone fights to get more added whenever it comes up.
The key takeaway, though is that none of this complicated mess applies to physical books! So there ya go 😁👍


If you own a book, you can legally make as many copies of it as you want. As long as you’re not distributing them, the courts treat that as a single effective copy.


Legal reasons: When you “purchase” an ebook you’re actually just licensing it and nearly all ebook licenses exclude the ability to do anything with the ebook other than read it yourself.
They could get a commercial license to get big ebook libraries like Anthropic did, but they can’t, really, because Anthropic paid for an exclusive license. Which means that if other AI companies want to compete, they sort of have to buy books in bulk and tear them apart to scan them.


If you think any more than 0.1% of these physical books would ever have ended up in antique bookstores, you’re dreaming.
Think about how many books out there are things like Donald Trump’s biography, or pointless drivel from non-experts, self-help books that tell people to down “essential oils”, old editions of programming books, or just plain shitty fiction that never sold much in the first place.
It’s ok to throw trash away! Really!


That’s like saying, “they had some failure modes from the synthetic data, so they should just obviously stop trying forever.”
They’ll just fix the edge cases and move on. Like any programming task.


Yes. That does make sense.
If you bought a book, scanned it—destroying it in the process—then read it on your computer, that would be completely acceptable.
Why is it wrong when a corporation does the same thing?
They’re not claiming ownership of the copyrights, just ownership of a copy. Which is how copyright works.


Big AI mostly switched to synthetic training data anyway. The books they’re digitizing are being used to gather knowledge, not writing styles or logic (mostly).
As in, when you ask ChatGPT how long some book is, it can just go check (if it’s in the database). It’s also useful if you ask about that book or about knowledge contained in that book. It’ll even reference books now (if you demand that in your prompt).
It’s not the same as earlier LLM tech which relied on scanned text to figure out how to respond to any given prompt (from a language standpoint). The “language” part of LLMs is a solved problem now (thanks to the synthetic training). At least for English 🤷
Oh gods. I leave them everywhere 🤷