Uncensored thinking model with character — coding, fiction, and long-form reasoning.

40B dense model expanded from Qwen3.6-27B, Heretic-abliterated, Deckard/PDK character-tuned, then Claude 4.6 Opus reasoning distilled. NEO-CODE dual-imatrix GGUFs target long context stability and near-BF16 quality at practical quant sizes. 256K context. Vision via mmproj.

Uncensored / Heretic Thinking Coder Creative writing Di-IMatrix MAX 256K context Vision + mmproj Apache-2.0
Warning from the author: this model has character. No nanny filters. Not SFW if you ask for NSFW. Suggested min quant: Q4_K_S / IQ3_S or higher.

At a glance

DavidAU · GGUF
40Bdense params
96layers · 1275 tensors
256Knative context
98.4%Q8_0 vs BF16
94%IQ4_XS vs BF16
83–84%IQ2_M vs BF16

Interactive example gallery

Sample generations from the model card · click a tab
Open on Hugging Face

NEO-CODE Di-IMatrix quants

pick by VRAM / quality
QuantSize~BF16Notes
IQ2_M13.8 GB83–84%lightest
IQ3_M17.1 GBsolid mid
IQ4_XS20.6 GB94%sweet spot
IQ4_NL21.6 GB~94%balanced
Q4_K_S21.2 GBmin recommended
Q4_K_M22.6 GBgeneral default
Q5_K_S / M25–26 GBtools / quality
Q6_K30.2 GB~97%high fidelity
Q8_0 HIGH39.9 GB98.4%max GGUF

For images, also download one mmproj-*.gguf into the same folder as the main GGUF.

Suggested settings

from model card
Thinking · general

temp 1.0 · top_p 0.95 · top_k 20 · rep_pen 1.0

Thinking · coding

temp 0.6 · top_p 0.95 · top_k 20 · rep_pen 1.0

Instruct / no-think

temp 0.7 · top_p 0.80 · top_k 20 · presence_penalty 1.5

Creative 40B

temp ~0.7 · context ≥ 8–16K · lower quants: rep_pen 1.05–1.1

System prompt

Optional. Even Be vivid and precise. helps lower quants and reduces looping.

Run locally

llama.cpp · Ollama · LM Studio
Copied