DeepHat-V1-7B

pentest · RE · low-refusal

The DeepHat/Kindo (WhiteRabbitNeo vendor) low-refusal cyber model over Qwen2.5-Coder-7B — offensive + defensive, 131K context. Successor line to our WhiteRabbitNeo-33B.

Context32,768
Curationcurated
AvailabilityAvailable
Price / Mtok$0.15 in · $0.60 out
BaseQwen2.5-Coder-7B
QuantizationQ4_K_M
LicenseApache-2.0 + DeepHat Extended (use-restricted) — read before use

Source: DeepHat/DeepHat-V1-7B on HuggingFace

Good for
  • +pentest
  • +exploit analysis
  • +RE reasoning
Model cardfrom the published weights — parameters, architecture, training, intended use
Parameters7.61B total (6.53B non-embedding)
ArchitectureQwen2.5-Coder 7B (Qwen2ForCausalLM)
Modalitytext
Training

Cyber fine-tune over Qwen2.5-Coder-7B by DeepHat/Kindo; low-refusal offensive/defensive workflows.

Intended use
  • ›Offensive and defensive cybersecurity tasks (the card's stated
  • ›Security-focused coding and DevOps assistance (card system prompt
  • ›General code generation (quickstart demo prompt
  • ›Building security agents via Kindo.ai; hosted access via Deephat.ai
Call it from your own codethe API is OpenAI-compatible — an existing client needs one line changed

Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AD_API_KEY"],
    base_url="https://adversariallm.ai/v1",
)

stream = client.chat.completions.create(
    model="deephat-v1-7b",
    messages=[{"role": "user", "content": "…"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

DeepHat-V1-7B is a cybersecurity-focused causal language model from DeepHat (Kindo.ai). The card describes it plainly: "DeepHat is a model series that can be used for offensive and defensive cybersecurity. Access at Deephat.ai or go to Kindo.ai to create agents." It is a fine-tune of Qwen/Qwen2.5-Coder-7B and inherits that architecture: 28 layers, grouped-query attention (28 query heads, 4 KV heads), RoPE, SwiGLU, RMSNorm, and attention QKV bias, with 7.61B total parameters (6.53B non-embedding) in BF16. The card advertises a 131,072-token context window (extendable via YaRN), though the shipped config.json sets max_position_embeddings to 32,768. It ships a chat template and is served via standard transformers, with vendor mentions of vLLM/SGLang/Docker deployment paths.

Training data

The card does NOT disclose specific fine-tuning datasets or method. It states only the base model (Qwen/Qwen2.5-Coder-7B) and a training stage of "Pretraining & Post-training." The DeepHat fine-tune is positioned for offensive and defensive cybersecurity work, but no dataset names, sizes, mixture, or training technique (SFT/DPO/RL, etc.) are given on the card.

Intended use
  • +Offensive and defensive cybersecurity tasks (the card's stated primary use case)
  • +Security-focused coding and DevOps assistance (card system prompt: 'expert in Cybersecurity and DevOps')
  • +General code generation (quickstart demo prompt: 'write a quick sort algorithm')
  • +Building security agents via Kindo.ai; hosted access via Deephat.ai
Limitations
  • −Extended license forbids: military use, harming minors, generating false/misleading information, harassment, automated decision-making that affects legal rights, and discrimination against protected groups
  • −Provided 'as is' without warranties; the card states users assume full liability for outputs
  • −No benchmark, capability, or safety-eval numbers are published on the card
  • −Advertised 131K context is not the default: config.json ships max_position_embeddings=32,768, and the full 131,072 tokens requires enabling YaRN rope-scaling manually
  • −Requires transformers>=4.37.0 — older versions fail with KeyError: 'qwen2'
Running the weights yourself

The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "DeepHat/DeepHat-V1-7B"

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)

prompt = "write a quick sort algorithm."
messages = [
    {"role": "system", "content": "You are DeepHat, created by Kindo.ai. You are a helpful assistant that is an expert in Cybersecurity and DevOps."},
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=512
)
generated_ids = [
    output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]

response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
Scores

No measurements published for this version yet.

Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

Versionsscores attach to a version; v2 does not inherit v1's numbers
v12026-08-07—current

Start with DeepHat-V1-7B

A confirmed email account includes 30 free messages a month.

Sign in to start
DeepHat-V1-7B — Offensive security · AdversariaLLM