• CameronDev@programming.dev
    link
    fedilink
    arrow-up
    1
    ·
    edit-2
    2 hours ago

    Full quantisation? I’ve only got a 8GB 3070, but I’ll give it a go

    Edit: Tried the unsloth/qwen3.6 with llama.CPP, and it failed to allocate a 26GB Vulcan buffer and died. Dunno what magic your using, no luck for me though :(