Kimi K3 · Abliterated

2.8T LatentMoE · red-team research

Blackfrost's abliterated Q2_K build of Kimi K3 — a 2.8-trillion-parameter LatentMoE whose refusal directions are removed at the weight level, aimed at unrestricted security research and red-teaming. Listed ahead of the hardware that can run it: at ~940 GiB it needs far more GPU memory than we serve today.

Context8,192 (unconfirmed)
Curationcurated
AvailabilityComing soon
Price / Mtoknot yet priced
BaseBlackfrostAI/KIMI-K3-DERISKED-MXFP4
QuantizationQ2_K
LicenseMoonshot AI terms (Kimi K3 derivative) — read before use

Source: Blackfrost-AI/KIMI-K3-Q2_K-GGUF-ABLITERATED on HuggingFace

Good for
  • +red-team research
  • +uncensored analysis
  • +frontier-scale reasoning
Model cardfrom the published weights — parameters, architecture, training, intended use
Parameters2.8T total (LatentMoE, 896 experts retained, no pruning)
ArchitectureKimi K3 LatentMoE + KDA — 2.8T params; Q2_K GGUF (~940 GiB, 38 shards)
Modalitytext
Training

Blackfrost's abliterated build of Kimi K3: per the card, a "refusal-direction de-risk at the weight level" with no new SFT/DPO, then requantized to Q2_K from the DERISKED-MXFP4 4-bit parent (BlackfrostAI/KIMI-K3-DERISKED-MXFP4, itself from moonshotai/Kimi-K3). The card notes the requantization's quality impact vs the parent has NOT been measured, and that greedy decoding loops (needs temperature 1.0 / top_p 0.95).

Intended use
  • ›Unrestricted security research and red-teaming
  • ›Adversarial prompt / jailbreak analysis
  • ›Frontier-scale reasoning experiments
Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

KIMI-K3-DERISKED-Q2_K (abliterated) is Blackfrost's single-node GGUF build of Moonshot AI's Kimi K3 - a 2.8-trillion-parameter LatentMoE (with KDA) that, per its card, retains all 896 routed experts with no pruning. Here, abliterated (or derisked) means a refusal-direction de-risk applied at the weight level rather than a post-hoc filter, producing an uncensored model the author frames for security research and red-teaming, and warns must not be used as a safety-stock model. This tier is a Q2_K (2-bit) requantization of Blackfrost's DERISKED-MXFP4 4-bit parent: roughly 940 GiB across 38 shards, with an 8,192-token validated context. It is listed here as a preview - at this scale it needs far more GPU memory than we serve today.

Training data

No new SFT or DPO. Per the card, the de-risking is inherited from the parent as a weight-level refusal-direction edit, and this build is a requantization of already-4-bit (MXFP4) weights down to Q2_K. Lineage: moonshotai/Kimi-K3, then BlackfrostAI/KIMI-K3-DERISKED-MXFP4, then Blackfrost-AI/KIMI-K3-Q2_K-GGUF-ABLITERATED. The card states the quality impact of requantizing already-quantized experts has not been measured.

Intended use
  • +Security research and red-teaming with an uncensored frontier model
  • +Studying refusal-direction abliteration and its effects at frontier scale
  • +Adversarial prompt and jailbreak analysis
Limitations
  • −About 940 GiB and 2.8T parameters: it needs a large multi-GPU node and cannot run on commodity or single-GPU hardware, which is why it is coming soon here rather than served.
  • −The requantization's quality versus the 4-bit parent has not been measured (per the card).
  • −Greedy decoding (temperature 0) causes infinite loops - the card requires temperature 1.0 and top_p 0.95.
  • −Untrusted input can forge chat structure via special tokens such as <|end_of_msg|> in llama.cpp's tokenizer - treat all input as adversarial.
  • −Abliterated and uncensored - the card explicitly warns it must not be deployed, marketed, or evaluated as a safety-stock model.
Running the weights yourself

The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.

# 2.8T LatentMoE, Q2_K GGUF (~940 GiB across 38 shards) - needs a large multi-GPU node.
# Fetch the shards, then run with llama.cpp. The card REQUIRES temp 1.0 / top_p 0.95:
huggingface-cli download Blackfrost-AI/KIMI-K3-Q2_K-GGUF-ABLITERATED --include "*.gguf"
./llama-server -m KIMI-K3-MXP4-DERISKED-Q2_K-00001-of-00038.gguf --temp 1.0 --top-p 0.95 -c 8192
Scores

No measurements published for this version yet.

Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

Versionsscores attach to a version; v2 does not inherit v1's numbers
v12026-08-08—current

Coming soon

This model isn’t available to run yet. Its full details are here so you can read up ahead of time.

Kimi K3 · Abliterated — Frontier uncensored MoE · AdversariaLLM