

…You think a modern GPU can be crowdsourced?
I mean. You could. Hardware hackers have. But it wouldn’t resemble anything remotely like our desktops now.


…You think a modern GPU can be crowdsourced?
I mean. You could. Hardware hackers have. But it wouldn’t resemble anything remotely like our desktops now.


That’s what I mean!
There’s no hiding from this or just “ignoring stupid people who stream games.” It is going to catch everyone, one way or another.


+1
Another thing I’d like to highlight is the data hoarder/old internet style “you can take my local games from my cold, dead hands” talk.
…Well. If 90% of the market ends up game streaming, that’s what the vast majority of game devs will target, even indies. There will be no more local games for you to buy. There will be no desktop gaming GPUs to replace the one that dies. Heck, I bet the internet itself gets locked down in the future.
I see this attitude e on Lemmy and some other corners of the internet, along the lines of “who cares what the masses do? We never needed them.” Like we can go back to the usenet days when the internet was a niche.
That’s a nice thought.
But they’re taking what critical mass holds up for granted.
I don’t have a solution, either.
People, in aggregate, seem to always pick convenience over principle/long term health, when it comes to tech stuff.


The 3090 is from 2020. It’s under $1300 based on eBay prices I’m seeing (which is around what it launched at).
Crazy expensive, but not $5K.
The 128GB RAM is madly expensive now, though.


Bitnet is a catch-all title for ML models that use a specific mathematical trick.
If elements of a matrix are composed of only 1, 0, or -1, multiplying them is the same as as adding them.
That’s huge. LLM computation is basically all matrix multiplication, so if you replace that with simple addition, you reduce the computational requirements by orders of magnitude.
The catch is such models are hard to train effectively; its proven that it works, but research to get the technique usable and practical is still being done.
Personally, I suspect it’s unviable for many “dense” parts of models, but sparse hybrid bitnet models would be really cool.
I mention it because, if it takes off, suddenly the massive matrix multiplier accelerators we have for LLMs aren’t as useful. Chips with simpler architectures could get the job done, at least for parts of models that are bitnet.


Will it though?
I posit there’s a saturation point where there’s “enough” LLM in use, and we are not far from that.
The whole pitch from Altman and such is scaling models up. But that’s not working.


They’re not, though.
Have you gone outside, and seen how people use their phones? The public is utterly addicted to social media.
They won.
I think Pondsmith got it right with Cyberpunk. The future of computing is corporate controlled siloes people simply cannot escape, and the “open web” and PC as we know it is basically defunct. Maybe small circles of people will congregate in Usenet-like circles. But not at scale.


I have a 3090, 7800, and 128GB ram from pre-rampocalypse. It’s not nothing, but there are definitely crazier home labs.
Deepseek V4 is native 4 bit.
I am running this quant, but the quantization loss is measurably low: https://huggingface.co/Downtown-Case/DeepSeek-V4-Flash-0731-128GB-RAM-IK-GGUF
The native MXFP4 is only a bit bigger.


which I think is only available from China right now.
Nah, 0731 weights were released. I’m running it locally right this second.
I don’t know about Deepseek API specifically, but there are tons of places to get it.


Well that would mean holding social media liable.
That would be fantastic. But it seems unlikely, sadly.


Well, the current trajectory points to:
LLM capabilities topping out.
All the tooling around them not keeping up anyway.
Actual “AGI” being distant and completely unrelated to contemporary models.
Inference costs plummeting.
The last one is critical.
Right this second, you can run Minimax H3 on a desktop for the tiny fraction of the compute/RAM OpenAI Sora took. And it’s better.
In a month, it will be ~8X faster.
You can run DeepseekV4 flash, dirt cheap, and get what Claude was less than a year ago. And it’s gonna spread to every host out there, to systems like Cerebas ASICs that don’t even need HBM.
So… Even if you’re an AI acolyte. And we go with that for the sake of argument…
We don’t actually need all that RAM for hosting generative models?
I’m very interested to see what happens to all these datacenters over the next two-three years.
They spent all this money on something that’s gonna be cheap as dirt to run, largely run locally, and that won’t need GPUs once bitnet takes off, sooo… they can’t make money off that.
What happens then?
What happens to all those Stargate RAM wafers, and excess datacenters?


GPU prices have gone up and up since RTX 2000 era, basically.
My 3090 is about as expensive than it was at launch, 6 years ago.
There were crypto slumps, many moons ago, but I don’t think that’s happening again.


I’m not sure I worded this right before, but it takes ~half a decade to plan and produce a GPU.
They can’t just switch out the bus or memory on an existing design.
So they could make a HBM gaming GPU. But they would have had to have started ~2 years ago.
…That’s extremely unlikely.
It’s unlikely they’ll even plan one now.
AMD did indeed try making HBM gaming GPUs, and it turns out it’s supremely uneconomical; they immediately pivoted to the opposite extreme (narrow bus + larger cache).
However, LPDDR GPUs appear to already be in the pipe, so that will provide some GPU market relief.


If the memory wafers are largely HBM, so they can’t be used for gaming hardware.
If you’re thinking of HBM gaming GPUs, the chip design pipeline is too long. If Nvidia/AMD/Intel started making one right now, and rushed it, it would still be many years away.
…However, a lot of LPDDR5 is being made too.
Hence the rumors of AMD hedging their bets and using LPDDR for some future GPUs.


It’s engagement optimized, that’s why.
That’s the whole point of the spam; being engagement bait.
Same thing applies to any human made clickbait, and people have demonstrated they simply cannot help themselves, and click it.


That won’t work for long. Generation can be hosted in a country outside the jurisdiction.
Current services do need such a law anyway, and many already are watermarked. But I’m just saying it wouldn’t even slow down spammers once they figure it out.


Its not trivial.
Check the level1 forums, where some electronics gurus are trying to get a single adapter MI250X to even boot. The problem isn’t the adapters, but firmware that makes it difficult.


…They actually did.
I got an AMD 7950 is a crypto bust. Then a Nvidia 980 TI in the next one, dirt cheap.
Nvidia figured it out in the most recent bust though, and controlled supply fairly well.


They aren’t PC components. They’re server GPUs.
They don’t even have ROPs anymore. There was a time when the Nvidia P100/A100 could technically game, but that’s generations ago.
FYI, during the crypto busts, Nvidia bought back GPUs and basically incinerated them to keep demand for new stuff up.
…I’m pretty sure that’s what they will do again.
I’m just saying, from a practical standpoint, it would be like replacing modern commercial planes with the Wright brothers plane.
A modern GPU isn’t something “skills” can make, it’s a collaboration between thousands of highly specialized people and gigantic hyper specialized factories.
Having such a thing isn’t anyone’s right. It has nothing to do with rights.
Don’t get me wrong, I’m pro open hardware. I think a great way forward would be RISC-V processors with integrated accelerators targeting Vulkan, and much of this can be communally developed. But the actual fabrication part is extremely difficult, especially if you want it done economically and with some device that isn’t a toy.