BeefAndPoultry@lemmus.org to LocalLLaMA@sh.itjust.worksEnglish · 1 month agounsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Facehuggingface.coexternal-linkmessage-square7linkfedilinkarrow-up10arrow-down10file-text
arrow-up10arrow-down1external-linkunsloth/DeepSeek-V4-Flash-0731-GGUF · Hugging Facehuggingface.coBeefAndPoultry@lemmus.org to LocalLLaMA@sh.itjust.worksEnglish · 1 month agomessage-square7linkfedilinkfile-text
minus-squareMultiplexer@discuss.tchncs.delinkfedilinkEnglisharrow-up0·1 month agoAnyone knows, why 4bit quant is only marginally smaller than 8bit quant, though? And shouldn’t 8bit be roughly equal to the number of parameters anyway, so ~300GB?
minus-squareBeefAndPoultry@lemmus.orgOPlinkfedilinkEnglisharrow-up0·1 month agothe original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher
Anyone knows, why 4bit quant is only marginally smaller than 8bit quant, though?
And shouldn’t 8bit be roughly equal to the number of parameters anyway, so ~300GB?
the original model has a lot of parts that were natively trained in 4 bit, so those layers can’t go higher