It’s definitely better than American open-weight models. The gap between American flagship models and Chinese open-weight models is closing quickly though.
With the token prices of Chinese models, I seldom reach for American models anymore.
I’ve spent the last week or so coding an application using nothing but Qwen3.8-27B. I was curious how much I could do entirely with a local model running on my own personal hardware, and it turns out the answer was “everything.”
Granted, it’s not the fanciest application. But it’s fun.
I’m really looking forward to seeing what Qwen4 is like. 3.8 was just 3.6 with additional training, my understanding is that Qwen4 is new from the ground up.
Sadly they don’t seem to have 36B-A3B models on their roadmap any more, they didn’t do one for 3.8 either. I agree it was a nice sweet spot between speed and capability, I still use the 3.6 version of 36B-A3B for larger-scale local work. Maybe someone else will aim for that. Or Qwen4 will do something new with the architecture that makes it unnecessary.
It’s definitely better than American open-weight models. The gap between American flagship models and Chinese open-weight models is closing quickly though.
With the token prices of Chinese models, I seldom reach for American models anymore.
I’ve spent the last week or so coding an application using nothing but Qwen3.8-27B. I was curious how much I could do entirely with a local model running on my own personal hardware, and it turns out the answer was “everything.”
Granted, it’s not the fanciest application. But it’s fun.
Yep, I’ve been using that same model pretty heavily too, running on ExLlamaV3. It’s very impressive for such a little model!
I’m really looking forward to seeing what Qwen4 is like. 3.8 was just 3.6 with additional training, my understanding is that Qwen4 is new from the ground up.
I hope they make a MoE version. I get like 50 tokens/sec with Qwen3.6-36B-A3B, but can’t rely on it quite as much as Qwen3.8-27B.
Sadly they don’t seem to have 36B-A3B models on their roadmap any more, they didn’t do one for 3.8 either. I agree it was a nice sweet spot between speed and capability, I still use the 3.6 version of 36B-A3B for larger-scale local work. Maybe someone else will aim for that. Or Qwen4 will do something new with the architecture that makes it unnecessary.
My Hermes instance periodically checks the overall open-weight LLM landscape, and it recently recommended trying Ornith 1.5 35B-A3B.
I haven’t tried it yet, but on paper, it sounds like it has some potential.