• Armand1@lemmy.world
    link
    fedilink
    English
    arrow-up
    5
    ·
    4 hours ago

    Having run models locally, RAM use seems to be almost directly proportional to number of parameters. 8 Billion parameters requires approx 8GB of VRAM at 1/4 precision.

    Therefore, if this pattern holds you somehow need 10 Terabytes of VRAM at 4K and 40 Terabytes at full precision.

    I think I saw some estimates that Claude’s Opus models may be and Opus model equivalents may be at around 100B parameters (100-400GB VRAM).

    TLDR its clear why RAM is so expensive.

    • Eager Eagle@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 hours ago

      for inference you’re only counting active parameters towards VRAM, and some labs / models don’t train at 32b precision, or even use the same precision for different parts of the network