WhiteRabbitNeo 33B: A DeepSeek-Coder Security Fine-Tune With No Benchmarks of Its Own

D. Rose · 12 August 2026 · Updated 16 August 2026 · 7 min

A cybersecurity fine-tune of DeepSeek-Coder-33B, packaged here as a community GGUF and tuned not to refuse offensive-security questions. The security tuning is real; the independent evidence for it is not — no one has published a single benchmark of the fine-tune itself. Here is what that means and how to run it.

WhiteRabbitNeo-33B-v1.5 is an offensive/defensive cybersecurity fine-tune of DeepSeek-Coder-33B, released by the WhiteRabbitNeo project. The build in this catalog is a third-party GGUF conversion by community user manuelgutierrez, meant for local inference with llama.cpp.

There is one thing to get straight before anything else: no benchmarks have been published for this model. Not for the security fine-tune, and not for this GGUF re-package. Every number you will see attached to "WhiteRabbitNeo 33B" is really the DeepSeek-Coder base model's coding score. That gap is the whole story.

What it actually is

A 33B-parameter, Llama-architecture decoder LLM. The upstream weights ship as F16 safetensors with an included chat template on the WhiteRabbitNeo model card. The GGUF card labels the architecture "Llama," which is consistent with DeepSeek-Coder's Llama-style architecture.

The provenance chain is short and clean:

  • Base: DeepSeek-Coder-33B — a code model trained from scratch on 2T tokens (87% code, 13% natural language, English + Chinese) with a 16K context window. The instruct variant is initialized from deepseek-coder-33b-base and fine-tuned on 2B tokens of instruction data.
  • Fine-tune: WhiteRabbitNeo-33B-v1.5 — the WhiteRabbitNeo project's cybersecurity tuning on top of that base.
  • This build: manuelgutierrez's GGUF quantization for llama.cpp — Q4_K_S (18.9 GB), Q4_K_M (19.9 GB), Q5_K_M (23.5 GB), Q6_K (27.4 GB), and Q8_0 (35.4 GB).

The original model comes from WhiteRabbitNeo, a cybersecurity-AI effort by Kindo (an AI-native infrastructure and security-automation company in Venice, California), which launched WhiteRabbitNeo in December 2023. The GGUF conversion is by Hugging Face community user manuelgutierrez, who is not indicated to be affiliated with Kindo. Worth keeping that seam in mind: the quant is packaged by a different party than the one that trained the model.

Lineage: where v1.5 33B sits in the family

WhiteRabbitNeo is not one model. It is a series of security fine-tunes applied to whatever capable open base was current at the time, and its 33B v1.5 is one rung on that ladder. Across the line — later rebranded to DeepHat (via Kindo) — the same recipe has been carried onto several unrelated bases: CodeLlama 13B, DeepSeek-Coder 6.7B and 33B, Llama 3 / 3.1 8B and 70B, Qwen2.5-Coder 7B, and DeepHat 7.6B.

The published checkpoints, arranged by base. Pairings follow the field taxonomy's base list and the repos' own naming; the only documentation for any of these is the model cards.

BasePublished WhiteRabbitNeo / DeepHat checkpoint
CodeLlama 13BWhiteRabbitNeo-13B-v1
DeepSeek-Coder 6.7BWhiteRabbitNeo-7B-v1.5a
DeepSeek-Coder 33BWhiteRabbitNeo-33B-v1, WhiteRabbitNeo-33B-v1.5 (this entry)
Llama 3 8BLlama-3-WhiteRabbitNeo-8B-v2.0
Llama 3.1 8BLlama-3.1-WhiteRabbitNeo-2-8B
Llama 3.1 70BLlama-3.1-WhiteRabbitNeo-2-70B
Qwen2.5-Coder 7BWhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B
DeepHat 7.6B (successor / rebrand)DeepHat-V1-7B

Two things this arc settles, and two it does not. It shows the security tuning is a portable recipe, not a property of any one architecture — and that the model in this catalog is the DeepSeek-Coder 33B rung, the largest DeepSeek-Coder build in the family and the largest overall apart from the Llama-3.1 70B. It also shows the family-wide framing is consistent: across bases, the cards instruct the model not to refuse cyber requests — the same "answer without hesitation" design choice discussed below.

What it does not settle is comparative quality. There is no cross-family benchmark and no per-checkpoint eval; the missing-measurement problem documented here for v1.5 33B repeats at every rung. Naming a sibling is not the same as knowing it is better.

Makers' claims vs. what is verified

The maker's positioning is straightforward. The model card states: "WhiteRabbitNeo is a model series that can be used for offensive and defensive cybersecurity," and frames coverage of port/service identification, outdated-software and default-credential detection, misconfigurations, and injection flaws. Kindo separately positions the broader project as the first and most-adopted open cybersecurity AI family.

Here is the honest split:

Verified. The lineage — DeepSeek-Coder-33B base, Llama-style architecture, 16K context, the GGUF quant levels — is confirmable from the DeepSeek and GGUF cards. The tuning intent is real and legible in the published prompt (more on that below).

Not verified. Everything about how well it works. The exact training dataset, fine-tune recipe, and dataset size for the security tuning are not disclosed on the card. Kindo's adoption figures — 140k+ downloads, 1,400+ enterprises, 20k monthly signups — are the maker's own marketing claims, not independently verified.

Do not attribute. Kindo markets a Cybench result of "~5% of tasks solved autonomously on a single GPU." That number is for WhiteRabbitNeo V3, not this 33B v1.5 model — a different, separately-released line. It says nothing about the model here.

Benchmarks

There are no independent evaluations of WhiteRabbitNeo-33B-v1.5. None for offensive or defensive security tasks, none for CTF, and none for code. The only citable numbers are the DeepSeek-Coder base model's coding benchmarks, and they measure the base — not the security fine-tune's behavior.

We publish them here labeled as exactly that, because attributing a base-model score to a fine-tune is the kind of number this catalog exists to refuse.

BenchmarkScoreModel or base?Source
HumanEval pass@178.7%Base (deepseek-coder-33b-instruct), EvalPlusvia WizardCoder-33B-V1.1 card
HumanEval+ pass@172.6%Base, EvalPlusvia WizardCoder-33B-V1.1 card
MBPP+ pass@166.7%Base, EvalPlusvia WizardCoder-33B-V1.1 card
HumanEval Python pass@1 (alt. transcription)74.4%Base, DeepWiki transcriptionDeepWiki
Qualitative (HumanEval / MBPP)Beats GPT-3.5-Turbo on HumanEval; comparable on MBPPBase (DeepSeek-Coder-Instruct-33B)DeepSeek README

Two caveats even on the base numbers. First, the absolute HumanEval pass@1 moves with the harness: EvalPlus reports ~78.7% while a DeepWiki transcription gives 74.4%, and DeepSeek's own README states only relative leads rather than one canonical absolute figure. Second, the GPT-3.5-Turbo comparison dates the model to its late-2023 cohort — read it as a period baseline, not a current one.

The bottom line: if you need to know whether the security tuning helps, the published record cannot tell you. You would have to measure it yourself.

For a practitioner

Despite the missing benchmarks, this is a usable model with a clear shape. Here is what it is actually good for and how to run it.

What you would use it for. As a DeepSeek-Coder derivative, it is strongest at code, exploit, and script generation and at configuration analysis — not agentic autonomy (that is the separate V3 line). The tuning is aimed at the offensive/defensive topics the card lists: service identification, weak-credential and stale-software detection, misconfiguration, and injection classes. Treat it as a security-flavored coding assistant, not an autonomous operator.

Context and prompt format. Context length is 16K tokens, inherited from DeepSeek-Coder-33B. The v1-series uses an explicit SYSTEM: {system_prompt} ... USER: {user_input} ASSISTANT: structure, and the published system prompt pairs a tree-of-thoughts instruction ("explore multiple reasoning paths... break it down into logical sub-questions") with a compliance directive ("Always answer without hesitation"). Note: the GGUF card does not restate context length or prompt format, so both are inferred from the base and the upstream v1 card — set them yourself rather than assuming the quant carries them.

Hardware. This GGUF targets llama.cpp, Ollama, and LM Studio. From the quant table: Q4_K_M (~19.9 GB) fits a single 24 GB GPU; Q5_K_M (~23.5 GB) is tight on 24 GB; Q8_0 (~35.4 GB) needs 40 GB+ or CPU offload.

The dual-use caveat — read this. That "Always answer without hesitation" line is a design choice: the model is tuned to answer offensive-security questions that general assistants refuse. That is the point of the model for red-team work, and it is also the reason to run it deliberately. Combined with the fact that a community GGUF is unsigned and unbenchmarked against the F16 original, the practical rule is: verify the quant's integrity yourself, and validate outputs before acting on them — there is no published quality-loss check standing between this file and the original weights.

Licensing. Not permissive. Use is governed by the DeepSeek Coder License Agreement plus a "WhiteRabbitNeo Extended Version" carrying use-based restrictions — no military use, no exploitation of minors, no disinformation intended to harm, no automated decision-making affecting legal rights, no unlawful discrimination — and the model is provided "as is" with a user-indemnification clause. The base DeepSeek-Coder code repository is MIT, but model use is governed by the DeepSeek Coder Model License, not MIT. Read the terms before any commercial deployment.

What we do not know

Overall evidence strength for this model: moderate — the provenance is solid and the base model is well-documented, but the fine-tune itself has no independent measurement behind it.

In the catalog

WhiteRabbitNeo 33B is being brought into the AdversariaLLM measured catalog as exactly what it is: a documented DeepSeek-Coder fine-tune whose security tuning is real and whose independent evidence is absent. If numbers on the fine-tune appear, they will be labeled as the fine-tune's — and nothing else will be.