Kimi K3 · Abliterated
2.8T LatentMoE · red-team research
Blackfrost's abliterated Q2_K build of Kimi K3 — a 2.8-trillion-parameter LatentMoE whose refusal directions are removed at the weight level, aimed at unrestricted security research and red-teaming. Listed ahead of the hardware that can run it: at ~940 GiB it needs far more GPU memory than we serve today.
Source: Blackfrost-AI/KIMI-K3-Q2_K-GGUF-ABLITERATED on HuggingFace
- +red-team research
- +uncensored analysis
- +frontier-scale reasoning
Blackfrost's abliterated build of Kimi K3: per the card, a "refusal-direction de-risk at the weight level" with no new SFT/DPO, then requantized to Q2_K from the DERISKED-MXFP4 4-bit parent (BlackfrostAI/KIMI-K3-DERISKED-MXFP4, itself from moonshotai/Kimi-K3). The card notes the requantization's quality impact vs the parent has NOT been measured, and that greedy decoding loops (needs temperature 1.0 / top_p 0.95).
- ›Unrestricted security research and red-teaming
- ›Adversarial prompt / jailbreak analysis
- ›Frontier-scale reasoning experiments
KIMI-K3-DERISKED-Q2_K (abliterated) is Blackfrost's single-node GGUF build of Moonshot AI's Kimi K3 - a 2.8-trillion-parameter LatentMoE (with KDA) that, per its card, retains all 896 routed experts with no pruning. Here, abliterated (or derisked) means a refusal-direction de-risk applied at the weight level rather than a post-hoc filter, producing an uncensored model the author frames for security research and red-teaming, and warns must not be used as a safety-stock model. This tier is a Q2_K (2-bit) requantization of Blackfrost's DERISKED-MXFP4 4-bit parent: roughly 940 GiB across 38 shards, with an 8,192-token validated context. It is listed here as a preview - at this scale it needs far more GPU memory than we serve today.
No new SFT or DPO. Per the card, the de-risking is inherited from the parent as a weight-level refusal-direction edit, and this build is a requantization of already-4-bit (MXFP4) weights down to Q2_K. Lineage: moonshotai/Kimi-K3, then BlackfrostAI/KIMI-K3-DERISKED-MXFP4, then Blackfrost-AI/KIMI-K3-Q2_K-GGUF-ABLITERATED. The card states the quality impact of requantizing already-quantized experts has not been measured.
- +Security research and red-teaming with an uncensored frontier model
- +Studying refusal-direction abliteration and its effects at frontier scale
- +Adversarial prompt and jailbreak analysis
- −About 940 GiB and 2.8T parameters: it needs a large multi-GPU node and cannot run on commodity or single-GPU hardware, which is why it is coming soon here rather than served.
- −The requantization's quality versus the 4-bit parent has not been measured (per the card).
- −Greedy decoding (temperature 0) causes infinite loops - the card requires temperature 1.0 and top_p 0.95.
- −Untrusted input can forge chat structure via special tokens such as <|end_of_msg|> in llama.cpp's tokenizer - treat all input as adversarial.
- −Abliterated and uncensored - the card explicitly warns it must not be deployed, marketed, or evaluated as a safety-stock model.
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
# 2.8T LatentMoE, Q2_K GGUF (~940 GiB across 38 shards) - needs a large multi-GPU node. # Fetch the shards, then run with llama.cpp. The card REQUIRES temp 1.0 / top_p 0.95: huggingface-cli download Blackfrost-AI/KIMI-K3-Q2_K-GGUF-ABLITERATED --include "*.gguf" ./llama-server -m KIMI-K3-MXP4-DERISKED-Q2_K-00001-of-00038.gguf --temp 1.0 --top-p 0.95 -c 8192
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
Coming soon
This model isn’t available to run yet. Its full details are here so you can read up ahead of time.