• Coriza@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    18 hours ago

    The other day I saw a video talking about this new innovation on LLM inference side of things where they keep some more used weights in RAM and others less used on disk. I always suspected from the sample code I stumbled upon on the IA world that should be extreme opportunities for optimizations. But I cannot stress it enough how dumb the LLM world is where the basics of implementing an LRU cache is pass of as some big innovation. Like any half competent comp-sci or comp-eng professional know about the basics of mitigating this basics bottlenecks like “the data does not fit on available RAM”, “The disk is slow”, etc.

    So is not surprising that now that it seems that the “powerfulness” of this LLMs is starting to plateau that we would start to see some improvement in performance/resource utilization and hence running costs.

    • wewbull@feddit.uk
      link
      fedilink
      English
      arrow-up
      0
      ·
      16 hours ago

      That works for “Mixture of Experts” models. These are basically models with distinct sets of weights and only a subset of them will be used on any particular query. The rest can sit on a disk.

      It doesn’t work for dense models, where every weight is used all the time. There’s nothing inactive so a cache has nothing to exploit.