• percent@infosec.pub
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    I hope they make a MoE version. I get like 50 tokens/sec with Qwen3.6-36B-A3B, but can’t rely on it quite as much as Qwen3.8-27B.

    • FaceDeer@fedia.io
      link
      fedilink
      arrow-up
      0
      ·
      1 day ago

      Sadly they don’t seem to have 36B-A3B models on their roadmap any more, they didn’t do one for 3.8 either. I agree it was a nice sweet spot between speed and capability, I still use the 3.6 version of 36B-A3B for larger-scale local work. Maybe someone else will aim for that. Or Qwen4 will do something new with the architecture that makes it unnecessary.

      • percent@infosec.pub
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        My Hermes instance periodically checks the overall open-weight LLM landscape, and it recently recommended trying Ornith 1.5 35B-A3B.

        I haven’t tried it yet, but on paper, it sounds like it has some potential.