• squaresinger@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    4 hours ago

    I can run models on my 3yo midrange smartphone. Gemma-4-E2B totally runs on there. Bonsai-8B too. But neither is really good for most tasks.

    Heck, “a basic model” even runs locally on the 8MB RAM of an ESP32-S3, but it’s utterly worthless at anything.

    If you get a bit more into self-hosting AI it quickly becomes obvious that for any actually useful real world tasks you need at least 24GB you can dedicate to the LLM alone, and if this is fully GPU vRAM, the performance is way, way higher than on CPU, even with an NPU.