unreachable.cloud
  • Communities
  • Create Post
  • Create Community
  • heart
    Support Lemmy
  • search
    Search
  • Login
  • Sign Up
beep@piefed.world to Technology@lemmy.worldEnglish · 20 hours ago

The price of AI is crashing faster than the rate of Moore's Law — intelligence costs are in freefall, outpacing comparative technologies like compute, DNA sequencing, and lithium batteries

epoch.ai

external-link
message-square
47
link
fedilink
1
external-link

The price of AI is crashing faster than the rate of Moore's Law — intelligence costs are in freefall, outpacing comparative technologies like compute, DNA sequencing, and lithium batteries

epoch.ai

beep@piefed.world to Technology@lemmy.worldEnglish · 20 hours ago
message-square
47
link
fedilink
The plunging price of thought
epoch.ai
external-link
Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, faster than electricity, compute, batteries or DNA sequencing ever fell.

cross-posted from: https://piefed.world/c/tech/p/1434965/the-price-of-ai-is-crashing-faster-than-the-rate-of-moore-s-law-report-suggests-intellig

alert-triangle
You must log in or register to comment.
  • humanspiral@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    4 hours ago

    The cost decline for a given level of performance does tend to slow over time

    Even in their tests, there are big drops in 2026 models, and recent ones.

    A much more comprehensive and easy test is to follow this benchmark suite (AAi). Its a mix of medium/hard benchmarks, with bias for agentic coding. (sorry for url dump) https://artificialanalysis.ai/?endpoints=openai_gpt-5-2-codex%2Cazure_kimi-k2-thinking%2Camazon-bedrock_qwen3-coder-480b-a35b-instruct%2Camazon-bedrock_qwen3-coder-30b-a3b-instruct%2Ctogetherai_minimax-m2-5_fp4%2Ctogetherai_glm-5_fp4%2Ctogetherai_qwen3-next-80b-a3b-reasoning%2Cgoogle_gemini-3-pro_ai-studio%2Cgoogle_glm-4-7%2Cmoonshot-ai_kimi-k2-thinking_turbo%2Cnovita_glm-5_fp8&models=mimo-v2-5-pro%2Cgpt-6-sol-low%2Cgemini-3-5-flash-minimal%2Cgpt-5-6-luna-low%2Cclaude-fable-5-1-low%2Ckimi-k2-6%2Ckimi-k2-6-non-reasoning%2Cmimo-v2-0206%2Cglm-5-3-flash%2Cgpt-6-luna-xhigh%2Cglm-4.5%2Cgpt-5-5%2Cclaude-opus-5-low%2Cmimo-v2-5-0424%2Cclaude-sonnet-5%2Cmimo-v2-6-pro%2Cclaude-opus-4-5-thinking%2Cminimax-m3%2Cgpt-6-astra%2Cclaude-opus-5-5%2Cqwen3-8-27b-non-reasoning%2Cgpt-6-luna-medium%2Cclaude-fable-5-1%2Cgpt-5-6-luna%2Cmimo-v2-5-pro-non-reasoning%2Cminimax-m2-7%2Ck2-horizon-mova-36b-a4b%2Cclaude-opus-4-6-adaptive%2Cgpt-5-6-luna-medium%2Cgemini-3-1-flash-lite-preview%2Cgpt-5-4-pro%2Cgrok-4-3-medium%2Cnvidia-nemotron-3-super-120b-a12b%2Cgpt-6-luna-non-reasoning%2Cqwen3-8-27b-medium%2Cgpt-5-5-medium%2Cdeepseek-v4-pro-0424-non-reasoning%2Cgpt-6-luna%2Cgemini-3-8-flash-medium%2Cnvidia-nemotron-3-nano-30b-a3b-reasoning%2Cgrok-4-5%2Cgemini-3-flash-reasoning%2Cgpt-6-luna-high%2Cqwen3-8-27b-low%2Cmimo-v2-flash%2Cdeepseek-v4-pro%2Cqwen3-6-35b-a3b%2Cgemini-3-8-flash%2Cclaude-4-5-sonnet-thinking%2Cllama-4-maverick%2Cgrok-4-3%2Cclaude-opus-4-8%2Cmuse-spark-1-3%2Cqwen3-8-flash-next%2Cgpt-6-astra-low%2Ckimi-k2-5%2Cgpt-5-4%2Cqwen3-8-27b%2Cqwen3-8-max%2Cgpt-5-5-high%2Cclaude-sonnet-5-5%2Chy3%2Cclaude-opus-5%2Cgpt-5-4-mini%2Ck2-horizon-375b-a23b%2Cgemini-3-1-pro-preview%2Cgpt-6-luna-low%2Cgpt-6-sol%2Cgrok-4-6%2Cclaude-4-1-opus-thinking%2Cglm-5-3%2Cgpt-6-sol-non-reasoning%2Cclaude-sonnet-5-non-reasoning%2Cgpt-5-6-sol%2Cgpt-5-6-sol-xhigh%2Cdeepseek-v4-1-flash%2Cgemini-3-7-flash-low%2Cclaude-opus-5-5-xhigh%2Cclaude-sonnet-4-6-adaptive%2Cglm-5-2-non-reasoning%2Cclaude-opus-5-5-medium%2Cgpt-oss-120b%2Cclaude-opus-5-5-high%2Cglm-5-2%2Ckimi-k3%2Cgrok-4-3-low%2Cdeepseek-v4-flash%2Cclaude-opus-5-medium%2Cgemini-4-argon&agents=claude-code-fable-5-1-max-with-fallback%2Ckimi-code-cli-kimi-k3%2Cgrok-build-grok-4-7-xhigh%2Cmuse-code-muse-spark-1-3-max%2Cantigravity-sdk-gemini-3-8-flash-high%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-6-xhigh%2Ccodex-deepseek-v4-pro-0813-max%2Ccodex-deepseek-v4-flash-0731-max%2Cmuse-code-muse-spark-1-3-xhigh&coding-agents=cost&model-creators=zai%2Cgoogle%2Calibaba%2Copenai%2Canthropic%2Cxai%2Cmeta%2Cmistral%2Cdeepseek%2Cstepfun%2Cthinking-machines%2Ckimi%2Cifm%2Cminimax%2Cxiaomi&releases=claude-opus-5-5%2Cclaude-fable-5-1%2Cgpt-6-astra%2Cgpt-5-6-luna%2Cmuse-spark-1-3%2Cgrok-4-7%2Cmimo-v2-6-pro%2Cglm-5-3%2Cgemini-3-8-flash%2Cdeepseek-v4-1-flash%2Cminimax-m3&capability-models=gemini-3-5-flash-lite%2Cstep-5%2Cinkling%2Cmimo-v2-6-flash%2Cglm-5-3-flash%2Cmimo-v2-6-pro%2Cminimax-m3%2Cgpt-6-astra%2Cclaude-opus-5-5%2Cclaude-fable-5-1%2Cgpt-5-6-luna%2Cmuse-glimmer%2Cgrok-4-7%2Cnvidia-nemotron-3-ultra-550b-a55b%2Cgpt-6-luna%2Cgemini-3-8-flash%2Cmuse-spark-1-3%2Cqwen3-8-27b%2Cqwen3-8-max%2Cclaude-sonnet-5-5%2Cclaude-opus-5%2Ck2-horizon-375b-a23b%2Cgpt-6-sol%2Cglm-5-3%2Cgpt-5-6-sol%2Cdeepseek-v4-1-flash%2Cmistral-medium-3-5%2Ckimi-k3%2Cqwen3-8-flash-next&cost=evaluation-breakdown

    Opus 4.8 (may 2026), by far best model at the time, gets equaled by deepseek 4 pro (july 31) and 4.1 flash (sept 10th) at less than 1/100th the cost. MiMo 2.6 pro (sept 21) is even cheaper and beats opus 4.8 scores by a wide margin. sonnet 4.6 max to gpt luna high is also a 99% drop in a short time for lower performance models.

  • Coriza@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    7 hours ago

    The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like “the data does not fit on available RAM”, “The disk is slow”, etc.

    So is not surprising that now that it seems that the “powerfulness” of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.

    • wewbull@feddit.uk
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 hours ago

      That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

      It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.

  • tharien@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    9 hours ago

    Moore’s Law hasn’t been applicable for years.

  • thedeadwalking4242@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    10 hours ago

    I forsee a future for LLMs which is where the only people who will be using it is where they HAVE to use it. When it’s the only reasonable tool for the job. And it will be expensive and unprofitable to run.

    It has some use cases, but they are few. The market can’t support what they’re people want LLMs to be

  • anon_8675309@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    10 hours ago

    Thanks in part to China who release a lot of research free.

  • Blaster M@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    10 hours ago

    All that R&D, and they need to squeeze efficiency out more at this stage. Faster and more specialized is better than big chingus do everything, and means we can hopefully stop spending all that money on making more datacenters.

    Personally, I don’t use the big chingus models at all… I much prefer local gen if I use it for things. Much better privacy that way, and I’m not throwing money at these jokers.

    • classic@fedia.io
      link
      fedilink
      arrow-up
      0
      ·
      6 hours ago

      Are there some local models that one can use that are not giving our data away and on par with chat GBT?

  • Treczoks@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    13 hours ago

    I wonder how they plan to match “prices are in free fall” to “the AI industry will have to make trillions a year in order not to go bust”.

    On the other hand, “prices in free fall” might be the answer they got from AI…

    • MoffKalast@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      11 hours ago

      At this point the cloud model firms are basically banking on making a superinteligence before anyone else and taking over the planet, otherwise they go bankrupt. I wish I was kidding.

      Nvidia wins either way though, local models, cloud models, shovels always sell. So they have that overvalued but still realistic bedrock to build houses of cards on.

  • scytale@piefed.zip
    link
    fedilink
    English
    arrow-up
    0
    ·
    14 hours ago

    Yeah, OpenAI released Astra just a couple of weeks ago, then around last week they released Sol 6, which is around half the price of Astra. So you have around the same level of “reasoning” at half the price. They will never turn a profit, at least for the next several years or so.

  • Dr. Bob@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    15 hours ago

    So revenue is falling. The only way the bubble grows is forcing this shit into even more places?

  • naught101@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    15 hours ago

    the price of thought

    Oh fuck off.

    • latetolemmy@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      7 hours ago

      No actual thought here in this comment

    • blackshirt@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      14 hours ago

      These fuckers would sell us air if they could.

      • nailingjello@piefed.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        13 hours ago

  • Chozo@fedia.io
    link
    fedilink
    arrow-up
    0
    ·
    15 hours ago

    That’s cool and all, but when can I buy RAM again?

    • Rothe@piefed.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      14 hours ago

      In 4 years or never. The latter probably being the most likely, since they are not just keeping RAM from you for AI purposes. They don’t want you to own you own hardware anymore, so they just simply stop manufacturing consumergrade hardware.

      • corsicanguppy@lemmy.ca
        link
        fedilink
        English
        arrow-up
        0
        ·
        9 hours ago

        They don’t want you to own [your] own hardware anymore

        Daily, it seems, I realize I’m one of the non-telepathic people: I can’t immediately know chip-maker company CEO motives and goals, and I just don’t see the conspiracy you do. What am I thinking right now?

      • Simon_Shitewood@lemmy.ml
        link
        fedilink
        English
        arrow-up
        0
        ·
        13 hours ago

        Probably not never - CXMT are trying to aggressively expand to fill the market now the major players have left, but it will still be a few years before prices really come down as a result.

  • pHr34kY@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    16 hours ago

    Kinda sus that the cost of electricity stops at 1973.

  • CompactFlax@discuss.tchncs.de
    link
    fedilink
    English
    arrow-up
    0
    ·
    16 hours ago

    “Look how low our cost of inference is”

    “Pay no attention to the marketing budget that exceeds Coca-Cola’s for a small fraction of their revenues”

    What technical, fundamental reason is there for the crash in price? The article just accepts the MSRP as fact. The cost of training can indeed be spread over time but it’s not spread across enough time. The inference cost doesn’t actually drop in reality.

    • pageflight@piefed.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      13 hours ago

      Let’s just quickly check https://isaiprofitable.com/ :

      nope

      Nope. The answer is still “they’re burning cash.” Only the companies making silicon are raking it in.

    • Grimy@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      14 hours ago

      There are constantly new techniques being developed to speed up inference or reduce model size post training. There’s about a hundred different levers to pull that play on inference cost, and some of them don’t have much of an impact on quality.

      • Jiral@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        14 hours ago

        Yet all major players are still incapable of covering their costs and have to resort to all sorts of nasty accounting tricks to make it apear otherwise. How so?

        • Grimy@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          14 hours ago

          Because of the massive investments into datacenters. Why is everyone so quick to drink the kool-aid. They are spending stupid amounts of money to capture the market and force everyone into a subscription service, not because it cost money to actually run the models. They are all switching their models to fine tuned quants after the first two weeks as well.

          The open source community has found dozens of ways to lower costs, do you really think these big companies aren’t using the same tricks and haven’t developed even more advanced techniques?

          • Jiral@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            13 hours ago

            It doesn’t matter what tricks they supposedly use afterwards, the data centets that are built need to run successfully or the bubble pops and then those companies will fall like lead. In other words for those data centers to be successfull the entire real economy has to spend crazy amounts of money on stuff that needs crazy amounts of compute.

            Open AI and Anthropic would really need some good business numbers, right now. The claim that they are just deliberately making their business case much worse than it is, is insane.

            • Grimy@lemmy.world
              link
              fedilink
              English
              arrow-up
              0
              ·
              13 hours ago

              The inference cost doesn’t actually drop in reality.

              My comment was about this which is just completely wrong.

              OpenAI is the scummiest company on earth but everyone is ready to believe them when they say shit like "our 20$ plan lets you run up the equivalent of a billion dollars in API fees. Giggles. It’s a good plan for you but not for us ;) "

              The main thing bringing up costs is Sam Altman spending all the “inference” money on raw wafers just to strangle consumer GPU sales.

  • Th4tGuyII@fedia.io
    link
    fedilink
    arrow-up
    0
    ·
    17 hours ago

    So in essence the price that the market will bare for the cost of AI usage is significantly lower than what the big AI companies would like it to be (in order to pay back their ever growing debts), which means there is a possibility they might never achieve profitability on their own (without some external/governmental strong-arming)

    • pageflight@piefed.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      13 hours ago

      Yeah, I’d like to see the cost of sub-prime mortgage on that chart.

  • Eq0@literature.cafe
    link
    fedilink
    English
    arrow-up
    0
    ·
    17 hours ago

    Extremely interesting… so the depreciation of older models is extreme, while new models are constantly presented. And all the while no AI company is making any profit. This whole story is bonkers

    • bcgm3@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      12 hours ago

      Hadn’t you heard? Some day, one of these things is gonna cure cancer. We all just gotta kill ourselves subsidizing it until then!

    • sorter_plainview@lemmy.today
      link
      fedilink
      English
      arrow-up
      0
      ·
      13 hours ago

      The catch is

      the cost of a given level of AI performance

      And more interestingly the article itself says this

      And to this extent, when comparing price drops for AI to drops for other technologies for which we have price series, we are comparing apples and oranges.

      Even then, they decided to make it the headline. This is just like LLM bros doing things they don’t know anything about. Absolute garbage.

Technology@lemmy.world

technology@lemmy.world

Subscribe from Remote Instance

Create a post
You are not logged in. However you can subscribe from another Fediverse account, for example Lemmy or Mastodon. To do this, paste the following into the search field of your instance: [email protected]

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


  • @[email protected]
  • @[email protected]
  • @[email protected]
  • @[email protected]
Visibility: Public
globe

This community can be federated to other instances and be posted/commented in by their users.

  • 543 users / day
  • 2.03K users / week
  • 4.5K users / month
  • 5K users / 6 months
  • 0 local subscribers
  • 88.4K subscribers
  • 1.41K Posts
  • 23.8K Comments
  • Modlog
  • mods:
  • L3s@lemmy.world
  • enu@lemmy.world
  • Technopagan@lemmy.world
  • L4sBot@lemmy.worldB
  • BE: 0.19.20
  • Modlog
  • Instances
  • Docs
  • Code
  • join-lemmy.org