40B dense model expanded from Qwen3.6-27B, Heretic-abliterated, Deckard/PDK character-tuned, then Claude 4.6 Opus reasoning distilled. NEO-CODE dual-imatrix GGUFs target long context stability and near-BF16 quality at practical quant sizes. 256K context. Vision via mmproj.
| Quant | Size | ~BF16 | Notes |
|---|---|---|---|
| IQ2_M | 13.8 GB | 83–84% | lightest |
| IQ3_M | 17.1 GB | — | solid mid |
| IQ4_XS | 20.6 GB | 94% | sweet spot |
| IQ4_NL | 21.6 GB | ~94% | balanced |
| Q4_K_S | 21.2 GB | — | min recommended |
| Q4_K_M | 22.6 GB | — | general default |
| Q5_K_S / M | 25–26 GB | — | tools / quality |
| Q6_K | 30.2 GB | ~97% | high fidelity |
| Q8_0 HIGH | 39.9 GB | 98.4% | max GGUF |
For images, also download one mmproj-*.gguf into the same folder as the main GGUF.
temp 1.0 · top_p 0.95 · top_k 20 · rep_pen 1.0
temp 0.6 · top_p 0.95 · top_k 20 · rep_pen 1.0
temp 0.7 · top_p 0.80 · top_k 20 · presence_penalty 1.5
temp ~0.7 · context ≥ 8–16K · lower quants: rep_pen 1.05–1.1
Optional. Even Be vivid and precise. helps lower quants and reduces looping.