• percent@infosec.pub
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 days ago

    Yep, I’ve been using that same model pretty heavily too, running on ExLlamaV3. It’s very impressive for such a little model!

    • FaceDeer@fedia.io
      link
      fedilink
      arrow-up
      0
      ·
      3 days ago

      I’m really looking forward to seeing what Qwen4 is like. 3.8 was just 3.6 with additional training, my understanding is that Qwen4 is new from the ground up.

      • percent@infosec.pub
        link
        fedilink
        English
        arrow-up
        0
        ·
        3 days ago

        I hope they make a MoE version. I get like 50 tokens/sec with Qwen3.6-36B-A3B, but can’t rely on it quite as much as Qwen3.8-27B.

        • FaceDeer@fedia.io
          link
          fedilink
          arrow-up
          0
          ·
          3 days ago

          Sadly they don’t seem to have 36B-A3B models on their roadmap any more, they didn’t do one for 3.8 either. I agree it was a nice sweet spot between speed and capability, I still use the 3.6 version of 36B-A3B for larger-scale local work. Maybe someone else will aim for that. Or Qwen4 will do something new with the architecture that makes it unnecessary.

          • percent@infosec.pub
            link
            fedilink
            English
            arrow-up
            0
            ·
            3 days ago

            My Hermes instance periodically checks the overall open-weight LLM landscape, and it recently recommended trying Ornith 1.5 35B-A3B.

            I haven’t tried it yet, but on paper, it sounds like it has some potential.