• brucethemoose@lemmy.world
    link
    fedilink
    English
    arrow-up
    7
    arrow-down
    5
    ·
    edit-2
    20 hours ago

    This isn’t the smart way, though.

    What the homelabbers do (at least before the RAM crisis) is buy Xeon/TR/EPYC boards on the cheap, and then run gaming GPUs for hybrid inference.

    This is what I do. I run MiMo 2.5 at 8-10t/s on a 7800X3D/RTX 3090/128GB CPU RAM, more with Dflash. That’s a 300B model: it’s not even in the same class as Qwen 27B, which is what the dev in OP’s article is trying to run.

    And this is small-time: setups with 4-8 memory channels can run stuff like Kimi or Deepseek Pro, even faster. Or they can run smaller LLMs with quantization types that are very fast on CPUs, and get crazy speeds.

    …And besides, Qwen 27B can run fine on a 4080, with the right framework. It will fit in 16GB as an exl3.


    Not that this isn’t a cool hardware hacking project.

    …But it’s kind of the wrong approach. It’s about 2 years out of date, as MoEs are king in LLM land now. RAM is horrendously expensive, yes, but so are most used V100s, or used 3090s.