5b3d1372b2466d64eb53b27f76f3b8783ac9f211
Model: ZeroWw/Meta-Llama-3.1-8B-Claude-39fail-3000total-GGUF Source: Original Platform
license, language, pipeline_tag
| license | language | pipeline_tag | |
|---|---|---|---|
| mit |
|
text-generation |
My own (ZeroWw) quantizations. output and embed tensors quantized to f16. all other tensors quantized to q5_k or q6_k.
Result: both f16.q6 and f16.q5 are smaller than q8_0 standard quantization and they perform as well as the pure f16.
Updated on: Wed Jul 24, 19:15:47
Description