• percent@infosec.pub
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 days ago

    It’s definitely better than American open-weight models. The gap between American flagship models and Chinese open-weight models is closing quickly though.

    With the token prices of Chinese models, I seldom reach for American models anymore.

    • FaceDeer@fedia.io
      link
      fedilink
      arrow-up
      0
      ·
      2 days ago

      I’ve spent the last week or so coding an application using nothing but Qwen3.8-27B. I was curious how much I could do entirely with a local model running on my own personal hardware, and it turns out the answer was “everything.”

      Granted, it’s not the fanciest application. But it’s fun.

      • percent@infosec.pub
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        Yep, I’ve been using that same model pretty heavily too, running on ExLlamaV3. It’s very impressive for such a little model!

        • FaceDeer@fedia.io
          link
          fedilink
          arrow-up
          0
          ·
          1 day ago

          I’m really looking forward to seeing what Qwen4 is like. 3.8 was just 3.6 with additional training, my understanding is that Qwen4 is new from the ground up.

          • percent@infosec.pub
            link
            fedilink
            English
            arrow-up
            0
            ·
            1 day ago

            I hope they make a MoE version. I get like 50 tokens/sec with Qwen3.6-36B-A3B, but can’t rely on it quite as much as Qwen3.8-27B.

            • FaceDeer@fedia.io
              link
              fedilink
              arrow-up
              0
              ·
              1 day ago

              Sadly they don’t seem to have 36B-A3B models on their roadmap any more, they didn’t do one for 3.8 either. I agree it was a nice sweet spot between speed and capability, I still use the 3.6 version of 36B-A3B for larger-scale local work. Maybe someone else will aim for that. Or Qwen4 will do something new with the architecture that makes it unnecessary.

              • percent@infosec.pub
                link
                fedilink
                English
                arrow-up
                0
                ·
                1 day ago

                My Hermes instance periodically checks the overall open-weight LLM landscape, and it recently recommended trying Ornith 1.5 35B-A3B.

                I haven’t tried it yet, but on paper, it sounds like it has some potential.