• BJW@lemmus.org
    link
    fedilink
    English
    arrow-up
    0
    ·
    10 hours ago

    Hosting an accurate AI locally typically takes 24GB of VRAM. Not having sufficient VRAM makes you reliant on remote models as a paid service. You pay many more times for tokens what you could pay for RAM just once for limitless tokens.

    • utopiah@lemmy.ml
      link
      fedilink
      arrow-up
      0
      ·
      7 hours ago

      Same, that’s not a normal usage. Spending a lot of time on Lemmy/Reddit/HackerNews tech community might lead to believe running a local model is normal but very few people do. When they do it’s typically on GPUs or Mac precisely for their unified memory architecture… but it’s not a normal setup even for gamers. Also typically it’s to tinker locally and learn from than actually productivity. I’m repeating here, I’m not saying it’s not an interesting use case for some, just not what most people do.