All the tooling around them not keeping up anyway.
Actual “AGI” being distant and completely unrelated to contemporary models.
Inference costs plummeting.
The last one is critical.
Right this second, you can run Minimax H3 on a desktop for the tiny fraction of the compute/RAM OpenAI Sora took. And it’s better.
In a month, it will be ~8X faster.
You can run DeepseekV4 flash, dirt cheap, and get what Claude was less than a year ago. And it’s gonna spread to every host out there, to systems like Cerebas ASICs that don’t even need HBM.
So… Even if you’re an AI acolyte. And we go with that for the sake of argument…
We don’t actually need all that RAM for hosting generative models?
I’m very interested to see what happens to all these datacenters over the next two-three years.
They spent all this money on something that’s gonna be cheap as dirt to run, largely run locally, and that won’t need GPUs once bitnet takes off, sooo… they can’t make money off that.
What happens then?
What happens to all those Stargate RAM wafers, and excess datacenters?
What happens to all those Stargate RAM wafers, and excess datacenters?
The bigger question is what happens to the world-economy when the largest players bottom out and the companies betting on them are massively downgraded because of their huge investments in a fantasy about eternal growth. Hope everyone is setting aside extra cash; it’s gonna make 2008 look like a walk in the park.
Traditionally during such busts the government starts outright confiscating bank accounts or otherwise enacts some fiscal policy that renders all savings (especially those in cash) completely worthless in a last bid to pass the buck onto the 99%.
Well, the current trajectory points to:
LLM capabilities topping out.
All the tooling around them not keeping up anyway.
Actual “AGI” being distant and completely unrelated to contemporary models.
Inference costs plummeting.
The last one is critical.
Right this second, you can run Minimax H3 on a desktop for the tiny fraction of the compute/RAM OpenAI Sora took. And it’s better.
In a month, it will be ~8X faster.
You can run DeepseekV4 flash, dirt cheap, and get what Claude was less than a year ago. And it’s gonna spread to every host out there, to systems like Cerebas ASICs that don’t even need HBM.
So… Even if you’re an AI acolyte. And we go with that for the sake of argument…
We don’t actually need all that RAM for hosting generative models?
I’m very interested to see what happens to all these datacenters over the next two-three years.
They spent all this money on something that’s gonna be cheap as dirt to run, largely run locally, and that won’t need GPUs once bitnet takes off, sooo… they can’t make money off that.
What happens then?
What happens to all those Stargate RAM wafers, and excess datacenters?
Lower inference costs will lead to more demand for RAM
Will it though?
I posit there’s a saturation point where there’s “enough” LLM in use, and we are not far from that.
The whole pitch from Altman and such is scaling models up. But that’s not working.
The bigger question is what happens to the world-economy when the largest players bottom out and the companies betting on them are massively downgraded because of their huge investments in a fantasy about eternal growth. Hope everyone is setting aside extra cash; it’s gonna make 2008 look like a walk in the park.
Traditionally during such busts the government starts outright confiscating bank accounts or otherwise enacts some fiscal policy that renders all savings (especially those in cash) completely worthless in a last bid to pass the buck onto the 99%.
Deepseek v4 Flash is insanely good and basically free. Especially the latest model which I think is only available from China right now.
Nah, 0731 weights were released. I’m running it locally right this second.
I don’t know about Deepseek API specifically, but there are tons of places to get it.
You are Mr. moneybags or you’re running 2 or 4bit quantized.
I have a 3090, 7800, and 128GB ram from pre-rampocalypse. It’s not nothing, but there are definitely crazier home labs.
Deepseek V4 is native 4 bit.
I am running this quant, but the quantization loss is measurably low: https://huggingface.co/Downtown-Case/DeepSeek-V4-Flash-0731-128GB-RAM-IK-GGUF
The native MXFP4 is only a bit bigger.