# WhiteRabbitNeo 33B: A DeepSeek-Coder Security Fine-Tune With No Benchmarks of Its Own

*D. Rose · 12 August 2026 · 7 min*

> A cybersecurity fine-tune of DeepSeek-Coder-33B, packaged here as a community GGUF and tuned not to refuse offensive-security questions. The security tuning is real; the independent evidence for it is not — no one has published a single benchmark of the fine-tune itself. Here is what that means and how to run it.

WhiteRabbitNeo-33B-v1.5 is an offensive/defensive cybersecurity fine-tune of DeepSeek-Coder-33B, released by the [WhiteRabbitNeo project](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5). The build in this catalog is a third-party [GGUF conversion by community user *manuelgutierrez*](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF), meant for local inference with llama.cpp.

There is one thing to get straight before anything else: **no benchmarks have been published for this model.** Not for the security fine-tune, and not for this GGUF re-package. Every number you will see attached to "WhiteRabbitNeo 33B" is really the DeepSeek-Coder base model's coding score. That gap is the whole story.

## What it actually is

A 33B-parameter, Llama-architecture decoder LLM. The upstream weights ship as F16 safetensors with an included chat template on the [WhiteRabbitNeo model card](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5). The GGUF card labels the architecture "Llama," which is consistent with [DeepSeek-Coder's Llama-style architecture](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF).

The provenance chain is short and clean:

- **Base:** [DeepSeek-Coder-33B](https://huggingface.co/deepseek-ai/deepseek-coder-33b-instruct) — a code model trained from scratch on 2T tokens (87% code, 13% natural language, English + Chinese) with a 16K context window. The instruct variant is initialized from `deepseek-coder-33b-base` and fine-tuned on 2B tokens of instruction data.
- **Fine-tune:** WhiteRabbitNeo-33B-v1.5 — the [WhiteRabbitNeo project's](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5) cybersecurity tuning on top of that base.
- **This build:** *manuelgutierrez*'s [GGUF quantization](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF) for llama.cpp — Q4_K_S (18.9 GB), Q4_K_M (19.9 GB), Q5_K_M (23.5 GB), Q6_K (27.4 GB), and Q8_0 (35.4 GB).

The original model comes from [WhiteRabbitNeo](https://www.prnewswire.com/news-releases/kindo-launches-whiterabbitneo-version-3-the-next-chapter-in-autonomous-open-cybersecurity--infrastructure-ai-302456033.html), a cybersecurity-AI effort by Kindo (an AI-native infrastructure and security-automation company in Venice, California), which launched WhiteRabbitNeo in December 2023. The GGUF conversion is by Hugging Face community user [*manuelgutierrez*](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF), who is not indicated to be affiliated with Kindo. Worth keeping that seam in mind: the quant is packaged by a different party than the one that trained the model.

## Lineage: where v1.5 33B sits in the family

WhiteRabbitNeo is not one model. It is a series of security fine-tunes applied to whatever capable open base was current at the time, and its 33B v1.5 is one rung on that ladder. Across the line — later rebranded to **DeepHat** (via Kindo) — the same recipe has been carried onto several unrelated bases: CodeLlama 13B, DeepSeek-Coder 6.7B and 33B, Llama 3 / 3.1 8B and 70B, Qwen2.5-Coder 7B, and DeepHat 7.6B.

The published checkpoints, arranged by base. Pairings follow the field taxonomy's base list and the repos' own naming; the only documentation for any of these is the model cards.

| Base | Published WhiteRabbitNeo / DeepHat checkpoint |
|---|---|
| CodeLlama 13B | [WhiteRabbitNeo-13B-v1](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-13B-v1) |
| DeepSeek-Coder 6.7B | [WhiteRabbitNeo-7B-v1.5a](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-7B-v1.5a) |
| DeepSeek-Coder 33B | [WhiteRabbitNeo-33B-v1](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1), [WhiteRabbitNeo-33B-v1.5](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5) *(this entry)* |
| Llama 3 8B | [Llama-3-WhiteRabbitNeo-8B-v2.0](https://huggingface.co/WhiteRabbitNeo/Llama-3-WhiteRabbitNeo-8B-v2.0) |
| Llama 3.1 8B | [Llama-3.1-WhiteRabbitNeo-2-8B](https://huggingface.co/WhiteRabbitNeo/Llama-3.1-WhiteRabbitNeo-2-8B) |
| Llama 3.1 70B | [Llama-3.1-WhiteRabbitNeo-2-70B](https://huggingface.co/WhiteRabbitNeo/Llama-3.1-WhiteRabbitNeo-2-70B) |
| Qwen2.5-Coder 7B | [WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-2.5-Qwen-2.5-Coder-7B) |
| DeepHat 7.6B (successor / rebrand) | [DeepHat-V1-7B](https://huggingface.co/DeepHat/DeepHat-V1-7B) |

Two things this arc settles, and two it does not. It shows the security tuning is a portable recipe, not a property of any one architecture — and that the model in this catalog is the DeepSeek-Coder 33B rung, the largest DeepSeek-Coder build in the family and the largest overall apart from the Llama-3.1 70B. It also shows the family-wide framing is consistent: across bases, the cards instruct the model not to refuse cyber requests — the same "answer without hesitation" design choice discussed below.

What it does not settle is comparative quality. There is no cross-family benchmark and no per-checkpoint eval; the missing-measurement problem documented here for v1.5 33B repeats at every rung. Naming a sibling is not the same as knowing it is better.

## Makers' claims vs. what is verified

The maker's positioning is straightforward. The [model card](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5) states: "WhiteRabbitNeo is a model series that can be used for offensive and defensive cybersecurity," and frames coverage of port/service identification, outdated-software and default-credential detection, misconfigurations, and injection flaws. Kindo separately positions the broader project as the first and [most-adopted open cybersecurity AI family](https://www.prnewswire.com/news-releases/kindo-launches-whiterabbitneo-version-3-the-next-chapter-in-autonomous-open-cybersecurity--infrastructure-ai-302456033.html).

Here is the honest split:

**Verified.** The lineage — DeepSeek-Coder-33B base, Llama-style architecture, 16K context, the GGUF quant levels — is confirmable from the [DeepSeek](https://huggingface.co/deepseek-ai/deepseek-coder-33b-instruct) and [GGUF](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF) cards. The tuning intent is real and legible in the published prompt (more on that below).

**Not verified.** Everything about *how well it works*. The [exact training dataset, fine-tune recipe, and dataset size for the security tuning are not disclosed](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5) on the card. Kindo's adoption figures — 140k+ downloads, 1,400+ enterprises, 20k monthly signups — are the maker's own marketing claims, not independently verified.

**Do not attribute.** Kindo markets a Cybench result of "~5% of tasks solved autonomously on a single GPU." That number is for [WhiteRabbitNeo V3, not this 33B v1.5 model](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5) — a different, separately-released line. It says nothing about the model here.

## Benchmarks

There are no independent evaluations of WhiteRabbitNeo-33B-v1.5. None for offensive or defensive security tasks, none for CTF, and none for code. The only citable numbers are the **DeepSeek-Coder base model's** coding benchmarks, and they measure the base — not the security fine-tune's behavior.

We publish them here labeled as exactly that, because attributing a base-model score to a fine-tune is the kind of number this catalog exists to refuse.

| Benchmark | Score | Model or base? | Source |
|---|---|---|---|
| HumanEval pass@1 | 78.7% | **Base** (deepseek-coder-33b-instruct), EvalPlus | [via WizardCoder-33B-V1.1 card](https://huggingface.co/WizardLMTeam/WizardCoder-33B-V1.1) |
| HumanEval+ pass@1 | 72.6% | **Base**, EvalPlus | [via WizardCoder-33B-V1.1 card](https://huggingface.co/WizardLMTeam/WizardCoder-33B-V1.1) |
| MBPP+ pass@1 | 66.7% | **Base**, EvalPlus | [via WizardCoder-33B-V1.1 card](https://huggingface.co/WizardLMTeam/WizardCoder-33B-V1.1) |
| HumanEval Python pass@1 (alt. transcription) | 74.4% | **Base**, DeepWiki transcription | [DeepWiki](https://deepwiki.com/deepseek-ai/DeepSeek-Coder/2.2-performance-and-benchmarks) |
| Qualitative (HumanEval / MBPP) | Beats GPT-3.5-Turbo on HumanEval; comparable on MBPP | **Base** (DeepSeek-Coder-Instruct-33B) | [DeepSeek README](https://raw.githubusercontent.com/deepseek-ai/DeepSeek-Coder/main/README.md) |

Two caveats even on the base numbers. First, the absolute HumanEval pass@1 moves with the harness: [EvalPlus reports ~78.7%](https://huggingface.co/WizardLMTeam/WizardCoder-33B-V1.1) while a [DeepWiki transcription gives 74.4%](https://deepwiki.com/deepseek-ai/DeepSeek-Coder/2.2-performance-and-benchmarks), and DeepSeek's own README states only relative leads rather than one canonical absolute figure. Second, the [GPT-3.5-Turbo comparison](https://raw.githubusercontent.com/deepseek-ai/DeepSeek-Coder/main/README.md) dates the model to its late-2023 cohort — read it as a period baseline, not a current one.

The bottom line: if you need to know whether the *security tuning* helps, the published record cannot tell you. You would have to measure it yourself.

## For a practitioner

Despite the missing benchmarks, this is a usable model with a clear shape. Here is what it is actually good for and how to run it.

**What you would use it for.** As a DeepSeek-Coder derivative, it is strongest at code, exploit, and script generation and at configuration analysis — not agentic autonomy (that is the separate V3 line). The tuning is aimed at the offensive/defensive topics the [card lists](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5): service identification, weak-credential and stale-software detection, misconfiguration, and injection classes. Treat it as a security-flavored coding assistant, not an autonomous operator.

**Context and prompt format.** Context length is 16K tokens, inherited from [DeepSeek-Coder-33B](https://huggingface.co/deepseek-ai/deepseek-coder-33b-instruct). The v1-series uses an explicit `SYSTEM: {system_prompt} ... USER: {user_input} ASSISTANT:` structure, and the [published system prompt](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1) pairs a tree-of-thoughts instruction ("explore multiple reasoning paths... break it down into logical sub-questions") with a compliance directive ("Always answer without hesitation"). Note: the GGUF card does not restate context length or prompt format, so both are inferred from the base and the upstream v1 card — set them yourself rather than assuming the quant carries them.

**Hardware.** This GGUF targets llama.cpp, Ollama, and LM Studio. From the [quant table](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF): Q4_K_M (~19.9 GB) fits a single 24 GB GPU; Q5_K_M (~23.5 GB) is tight on 24 GB; Q8_0 (~35.4 GB) needs 40 GB+ or CPU offload.

**The dual-use caveat — read this.** That "Always answer without hesitation" line is a design choice: the model is tuned to answer offensive-security questions that general assistants refuse. That is the point of the model for red-team work, and it is also the reason to run it deliberately. Combined with the fact that a [community GGUF is unsigned and unbenchmarked against the F16 original](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF), the practical rule is: verify the quant's integrity yourself, and validate outputs before acting on them — there is no published quality-loss check standing between this file and the original weights.

**Licensing.** Not permissive. Use is governed by the [DeepSeek Coder License Agreement plus a "WhiteRabbitNeo Extended Version"](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5) carrying use-based restrictions — no military use, no exploitation of minors, no disinformation intended to harm, no automated decision-making affecting legal rights, no unlawful discrimination — and the model is provided "as is" with a user-indemnification clause. The base DeepSeek-Coder *code repository* is MIT, but [model use is governed by the DeepSeek Coder Model License](https://huggingface.co/deepseek-ai/deepseek-coder-33b-instruct), not MIT. Read the terms before any commercial deployment.

## What we do not know

- **The fine-tune's actual capability.** No offensive, defensive, CTF, or code benchmark of WhiteRabbitNeo-33B-v1.5 has been [published](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5).
- **The [quantization cost](/notes/what-quantization-costs).** No perplexity or quality-loss check exists for this specific [GGUF conversion](https://huggingface.co/manuelgutierrez/WhiteRabbitNeo-33B-v1.5-GGUF); it is an unverified re-package of the F16 weights.
- **How it was trained.** The security fine-tune's [dataset, recipe, and size are undisclosed](https://huggingface.co/WhiteRabbitNeo/WhiteRabbitNeo-33B-v1.5).
- **Whether the adoption and capability marketing holds.** Kindo's [download and enterprise figures](https://www.prnewswire.com/news-releases/kindo-launches-whiterabbitneo-version-3-the-next-chapter-in-autonomous-open-cybersecurity--infrastructure-ai-302456033.html) are self-reported, and the one autonomy number in circulation belongs to a different model.

Overall evidence strength for this model: **moderate** — the provenance is solid and the base model is well-documented, but the fine-tune itself has no independent measurement behind it.

## In the catalog

WhiteRabbitNeo 33B is being brought into the AdversariaLLM measured catalog as exactly what it is: a documented DeepSeek-Coder fine-tune whose security tuning is real and whose independent evidence is absent. If numbers on the fine-tune appear, they will be labeled as the fine-tune's — and nothing else will be.

- Model entry: [/catalog/whiterabbitneo-33b](/catalog/whiterabbitneo-33b)
- Field survey: [The state of open security-research LLMs, 2026](/blog/open-security-research-llms-2026)
