• Eager Eagle@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    3 hours ago

    for inference you’re only counting active parameters towards VRAM, and some labs / models don’t train at 32b precision, or even use the same precision for different parts of the network