unreachable.cloud
  • Communities
  • Create Post
  • Create Community
  • heart
    Support Lemmy
  • search
    Search
  • Login
  • Sign Up
beep@piefed.world to Technology@lemmy.worldEnglish · 6 days ago

The price of AI is crashing faster than the rate of Moore's Law — intelligence costs are in freefall, outpacing comparative technologies like compute, DNA sequencing, and lithium batteries

epoch.ai

external-link
message-square
66
link
fedilink
1
external-link

The price of AI is crashing faster than the rate of Moore's Law — intelligence costs are in freefall, outpacing comparative technologies like compute, DNA sequencing, and lithium batteries

epoch.ai

beep@piefed.world to Technology@lemmy.worldEnglish · 6 days ago
message-square
66
link
fedilink
The plunging price of thought
epoch.ai
external-link
Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, faster than electricity, compute, batteries or DNA sequencing ever fell.

cross-posted from: https://piefed.world/c/tech/p/1434965/the-price-of-ai-is-crashing-faster-than-the-rate-of-moore-s-law-report-suggests-intellig

alert-triangle
You must log in or register to comment.
  • Kaligalis@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    5 days ago

    Yeah. Not a surprise.

    everyone has bet on AI getting good enough to fully replace humans fast. CEOs mandated use of AI. The results weren’t as great as expected. CEOs started limiting AI use to counter the token cost explosion after employees found out how to waste tokens fast.

    At the same time, China’s AI development is driven by the party instead of companies. They aren’t so much interested in money as they are in the strategic solution to the demographic problem caused by the one-child policy (which worked a bit too well too fast). The companies there plan on making money by providing the compute (they literally have the power and a two digit number of nucular GW under construction right now).
    So their models are almost as good as US ones and freely downloadable to run on whatever hardware you want. That naturally limits the longterm-achievable AI token prices to little more than the cost of just providing the raw compute.

    US AI companies are also in a cut-throat competition for customers right from the start. That obviously doesn’t help to keep prices high. Currently, they all burn money so fast that it’s hard for a human mind to comprehend.
    None of the US AI companies will survive the next decade if they don’t actually are the only one making AI actually able to fully replace human workers. They will all go bankrupt and might take the whole US economy with them.
    Nvidia will probably be the real winner if they don’t fuck this up somehow (not sure if that is even possible) because they are the ones selling the shovels in this gold rush.

    In two decades, it will be normal to have the capabilities of current frontier models running locally on your Chinese phone.

  • Kairos@lemmy.today
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    Oh yeah “intelligence” is cheap AF…

  • sheetzoos@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    As the price of AI decreases, it will become more widespread and pervasive.

    • Kairos@lemmy.today
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      AI is already free. These are the subscriptions that companies and consumers who already want to use it pay for.

  • humanspiral@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    The cost decline for a given level of performance does tend to slow over time

    Even in their tests, there are big drops in 2026 models, and recent ones.

    A much more comprehensive and easy test is to follow this benchmark suite (AAi). Its a mix of medium/hard benchmarks, with bias for agentic coding. (sorry for url dump) https://artificialanalysis.ai/?endpoints=openai_gpt-5-2-codex%2Cazure_kimi-k2-thinking%2Camazon-bedrock_qwen3-coder-480b-a35b-instruct%2Camazon-bedrock_qwen3-coder-30b-a3b-instruct%2Ctogetherai_minimax-m2-5_fp4%2Ctogetherai_glm-5_fp4%2Ctogetherai_qwen3-next-80b-a3b-reasoning%2Cgoogle_gemini-3-pro_ai-studio%2Cgoogle_glm-4-7%2Cmoonshot-ai_kimi-k2-thinking_turbo%2Cnovita_glm-5_fp8&models=mimo-v2-5-pro%2Cgpt-6-sol-low%2Cgemini-3-5-flash-minimal%2Cgpt-5-6-luna-low%2Cclaude-fable-5-1-low%2Ckimi-k2-6%2Ckimi-k2-6-non-reasoning%2Cmimo-v2-0206%2Cglm-5-3-flash%2Cgpt-6-luna-xhigh%2Cglm-4.5%2Cgpt-5-5%2Cclaude-opus-5-low%2Cmimo-v2-5-0424%2Cclaude-sonnet-5%2Cmimo-v2-6-pro%2Cclaude-opus-4-5-thinking%2Cminimax-m3%2Cgpt-6-astra%2Cclaude-opus-5-5%2Cqwen3-8-27b-non-reasoning%2Cgpt-6-luna-medium%2Cclaude-fable-5-1%2Cgpt-5-6-luna%2Cmimo-v2-5-pro-non-reasoning%2Cminimax-m2-7%2Ck2-horizon-mova-36b-a4b%2Cclaude-opus-4-6-adaptive%2Cgpt-5-6-luna-medium%2Cgemini-3-1-flash-lite-preview%2Cgpt-5-4-pro%2Cgrok-4-3-medium%2Cnvidia-nemotron-3-super-120b-a12b%2Cgpt-6-luna-non-reasoning%2Cqwen3-8-27b-medium%2Cgpt-5-5-medium%2Cdeepseek-v4-pro-0424-non-reasoning%2Cgpt-6-luna%2Cgemini-3-8-flash-medium%2Cnvidia-nemotron-3-nano-30b-a3b-reasoning%2Cgrok-4-5%2Cgemini-3-flash-reasoning%2Cgpt-6-luna-high%2Cqwen3-8-27b-low%2Cmimo-v2-flash%2Cdeepseek-v4-pro%2Cqwen3-6-35b-a3b%2Cgemini-3-8-flash%2Cclaude-4-5-sonnet-thinking%2Cllama-4-maverick%2Cgrok-4-3%2Cclaude-opus-4-8%2Cmuse-spark-1-3%2Cqwen3-8-flash-next%2Cgpt-6-astra-low%2Ckimi-k2-5%2Cgpt-5-4%2Cqwen3-8-27b%2Cqwen3-8-max%2Cgpt-5-5-high%2Cclaude-sonnet-5-5%2Chy3%2Cclaude-opus-5%2Cgpt-5-4-mini%2Ck2-horizon-375b-a23b%2Cgemini-3-1-pro-preview%2Cgpt-6-luna-low%2Cgpt-6-sol%2Cgrok-4-6%2Cclaude-4-1-opus-thinking%2Cglm-5-3%2Cgpt-6-sol-non-reasoning%2Cclaude-sonnet-5-non-reasoning%2Cgpt-5-6-sol%2Cgpt-5-6-sol-xhigh%2Cdeepseek-v4-1-flash%2Cgemini-3-7-flash-low%2Cclaude-opus-5-5-xhigh%2Cclaude-sonnet-4-6-adaptive%2Cglm-5-2-non-reasoning%2Cclaude-opus-5-5-medium%2Cgpt-oss-120b%2Cclaude-opus-5-5-high%2Cglm-5-2%2Ckimi-k3%2Cgrok-4-3-low%2Cdeepseek-v4-flash%2Cclaude-opus-5-medium%2Cgemini-4-argon&agents=claude-code-fable-5-1-max-with-fallback%2Ckimi-code-cli-kimi-k3%2Cgrok-build-grok-4-7-xhigh%2Cmuse-code-muse-spark-1-3-max%2Cantigravity-sdk-gemini-3-8-flash-high%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-6-xhigh%2Ccodex-deepseek-v4-pro-0813-max%2Ccodex-deepseek-v4-flash-0731-max%2Cmuse-code-muse-spark-1-3-xhigh&coding-agents=cost&model-creators=zai%2Cgoogle%2Calibaba%2Copenai%2Canthropic%2Cxai%2Cmeta%2Cmistral%2Cdeepseek%2Cstepfun%2Cthinking-machines%2Ckimi%2Cifm%2Cminimax%2Cxiaomi&releases=claude-opus-5-5%2Cclaude-fable-5-1%2Cgpt-6-astra%2Cgpt-5-6-luna%2Cmuse-spark-1-3%2Cgrok-4-7%2Cmimo-v2-6-pro%2Cglm-5-3%2Cgemini-3-8-flash%2Cdeepseek-v4-1-flash%2Cminimax-m3&capability-models=gemini-3-5-flash-lite%2Cstep-5%2Cinkling%2Cmimo-v2-6-flash%2Cglm-5-3-flash%2Cmimo-v2-6-pro%2Cminimax-m3%2Cgpt-6-astra%2Cclaude-opus-5-5%2Cclaude-fable-5-1%2Cgpt-5-6-luna%2Cmuse-glimmer%2Cgrok-4-7%2Cnvidia-nemotron-3-ultra-550b-a55b%2Cgpt-6-luna%2Cgemini-3-8-flash%2Cmuse-spark-1-3%2Cqwen3-8-27b%2Cqwen3-8-max%2Cclaude-sonnet-5-5%2Cclaude-opus-5%2Ck2-horizon-375b-a23b%2Cgpt-6-sol%2Cglm-5-3%2Cgpt-5-6-sol%2Cdeepseek-v4-1-flash%2Cmistral-medium-3-5%2Ckimi-k3%2Cqwen3-8-flash-next&cost=evaluation-breakdown

    Opus 4.8 (may 2026), by far best model at the time, gets equaled by deepseek 4 pro (july 31) and 4.1 flash (sept 10th) at less than 1/100th the cost. MiMo 2.6 pro (sept 21) is even cheaper and beats opus 4.8 scores by a wide margin. sonnet 4.6 max to gpt luna high is also a 99% drop in a short time for lower performance models.

  • Coriza@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like “the data does not fit on available RAM”, “The disk is slow”, etc.

    So is not surprising that now that it seems that the “powerfulness” of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.

    • KingRandomGuy@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      5 days ago

      Keep in mind many of the optimizations you’re talking about (LRU caching of experts, for example) are only really relevant at the single user local inference scale. As in, an individual wants to run a big model on their machine, but they don’t have enough VRAM to fit the model and KV cache. Accordingly, you’re basically only describing hobbyist and research projects, which aren’t really representative of the AI inference industry as a whole.

      Commercial inference keeps everything resident in VRAM, so expert caching isn’t necessary. So these things won’t help costs. A lot of other low hanging fruit (like hierarchical KV cache) has also existed for a long time for production-ready inference engines.

      • Coriza@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        5 days ago

        I am not sure it would not help commercial solutions, if all experts are used all the time, sure, but if for example the usage is biased for some experts it would enable one machine to serve more users in parallel or save on VRAM or DRAM without compromising response time, hence cutting costs.

    • wewbull@feddit.uk
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

      It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.

  • tharien@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    Moore’s Law hasn’t been applicable for years.

    • Cort@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      Yeah gimme some more of that 14nm++++++

  • cinoreus@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    That’s cause you burnt money equivalent of gdp of Spain in that time frame, and I am being generous here.

  • thedeadwalking4242@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    I forsee a future for LLMs which is where the only people who will be using it is where they HAVE to use it. When it’s the only reasonable tool for the job. And it will be expensive and unprofitable to run.

    It has some use cases, but they are few. The market can’t support what they’re people want LLMs to be

  • anon_8675309@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    Thanks in part to China who release a lot of research free.

  • Blaster M@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    All that R&D, and they need to squeeze efficiency out more at this stage. Faster and more specialized is better than big chingus do everything, and means we can hopefully stop spending all that money on making more datacenters.

    Personally, I don’t use the big chingus models at all… I much prefer local gen if I use it for things. Much better privacy that way, and I’m not throwing money at these jokers.

    • classic@fedia.io
      link
      fedilink
      arrow-up
      0
      ·
      6 days ago

      Are there some local models that one can use that are not giving our data away and on par with chat GBT?

      • Blaster M@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        6 days ago

        Ollama, Open WebUI, and Qwen, or one of the edits of it. The latest Qwen 26B is in a very good delta of smart enough to know how to help with stuff (I use it as a second opinion/initial proofread code checker) and small enough to run on a macbook or a gaming pc, and Open WebUI is the self host ChatGPT-like interface that works with Ollama, the turn key ai model runner backend. Couple that with your own searxng instance self hosted to give it the ability to search the web for answers if you want to give it that. Run it all in a virtual machine, of course.

        For media content, wan2gp (wan2 for the gpu poor) has an all in one stop shop for self hosted audio, image, and video. Just beware people don’t like generated stuff in prod, but its useful for quick concepts, or to do certain types of edits of existing stuff.

        There’s also tts and speech recognition via the whisper models, which sounds a lot better than the speak and spell voices from the past (unless you enjoy the degeneracy of spamming JOHN MADDEN JOHN MADDEN in the Steven Hawking voice) and can transcribe words a lot faster (realtime) and more accurately than in the past. The best ones run via Vulkan and thus work on any card that supports Vulkan.

        If you run Nextcloud, you can also connect Ollama to it to get the Nextcloud AI working.

        • classic@fedia.io
          link
          fedilink
          arrow-up
          0
          ·
          5 days ago

          Great input, thank you I appreciate it. Next off for me is learning how to do a virtual machine! I do have ollama currently on a computer connected to a speech to text app. And I find that very useful to have

          • Blaster M@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            5 days ago

            Virt-manager for a gui, qemu-kvm and libvirtd for the hypervisor. You will need to pass thru a graphics card, which makes it unavailable for the host machine entirely. I use debian stable for the vm, no gui, just ssh to manage it.

  • Treczoks@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    I wonder how they plan to match “prices are in free fall” to “the AI industry will have to make trillions a year in order not to go bust”.

    On the other hand, “prices in free fall” might be the answer they got from AI…

    • MoffKalast@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      At this point the cloud model firms are basically banking on making a superinteligence before anyone else and taking over the planet, otherwise they go bankrupt. I wish I was kidding.

      Nvidia wins either way though, local models, cloud models, shovels always sell. So they have that overvalued but still realistic bedrock to build houses of cards on.

      • Quazatron@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        6 days ago

        Nvidia wins unless some other company starts selling cheaper, faster, more efficient matrix multiplication machines.

        I’ve read some articles about radically different inference architectures that may tilt the scales, but I know this is wishful thinking because I really would like Nvidia to fail badly.

        Linus_nvidia.gif

        • MoffKalast@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          5 days ago

          Plenty have tried, all have fallen over flat on their face when it comes to actually providing usable drivers or production at scale. AMD’s still completely half assing it even today and Intel’s OneAPI and Vino is a bloated joke.

          But yes I would love a future where AMD finally hires an actual software team.

  • scytale@piefed.zip
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    Yeah, OpenAI released Astra just a couple of weeks ago, then around last week they released Sol 6, which is around half the price of Astra. So you have around the same level of “reasoning” at half the price. They will never turn a profit, at least for the next several years or so.

  • Dr. Bob@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    So revenue is falling. The only way the bubble grows is forcing this shit into even more places?

  • naught101@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    the price of thought

    Oh fuck off.

    • latetolemmy@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      No actual thought here in this comment

    • blackshirt@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      These fuckers would sell us air if they could.

      • Agent641@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        5 days ago

        I saw it in a documentary called Total Recall

        • W98BSoD@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          5 days ago

      • nailingjello@piefed.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        6 days ago

  • Chozo@fedia.io
    link
    fedilink
    arrow-up
    0
    ·
    6 days ago

    That’s cool and all, but when can I buy RAM again?

    • Kaligalis@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      5 days ago

      Never if you are in the US. In a few years if you are allowed to buy Chinese.

    • Rothe@piefed.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      In 4 years or never. The latter probably being the most likely, since they are not just keeping RAM from you for AI purposes. They don’t want you to own you own hardware anymore, so they just simply stop manufacturing consumergrade hardware.

      • corsicanguppy@lemmy.ca
        link
        fedilink
        English
        arrow-up
        0
        ·
        6 days ago

        They don’t want you to own [your] own hardware anymore

        Daily, it seems, I realize I’m one of the non-telepathic people: I can’t immediately know chip-maker company CEO motives and goals, and I just don’t see the conspiracy you do. What am I thinking right now?

      • Simon_Shitewood@lemmy.ml
        link
        fedilink
        English
        arrow-up
        0
        ·
        6 days ago

        Probably not never - CXMT are trying to aggressively expand to fill the market now the major players have left, but it will still be a few years before prices really come down as a result.

Technology@lemmy.world

technology@lemmy.world

Subscribe from Remote Instance

Create a post
You are not logged in. However you can subscribe from another Fediverse account, for example Lemmy or Mastodon. To do this, paste the following into the search field of your instance: [email protected]

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


  • @[email protected]
  • @[email protected]
  • @[email protected]
  • @[email protected]
Visibility: Public
globe

This community can be federated to other instances and be posted/commented in by their users.

  • 601 users / day
  • 2.1K users / week
  • 4.48K users / month
  • 5.37K users / 6 months
  • 0 local subscribers
  • 88.6K subscribers
  • 1.57K Posts
  • 27.9K Comments
  • Modlog
  • mods:
  • L3s@lemmy.world
  • enu@lemmy.world
  • Technopagan@lemmy.world
  • L4sBot@lemmy.worldB
  • BE: 0.19.20
  • Modlog
  • Instances
  • Docs
  • Code
  • join-lemmy.org