BugTraceAI Apex-G4 26B: an abliterated Gemma 4 MoE with a compliance score, not a capability score
D. Rose · 16 August 2026 · 6 min
A GGUF, DPO-fine-tuned, deliberately uncensored offensive-security build of Google's Gemma 4 26B MoE, shipped by an anonymous author. The one performance number attached to it measures whether the model answers — not whether its exploits work.
BugTraceAI Apex-G4 26B is a fine-tune of Google's Gemma 4 26B A4B, packaged as a GGUF and marketed for offensive research. It ships exactly one performance number, and that number measures whether the model answers a request — not whether what it produces compiles, evades, or works. Here is what is actually in the box, what is only claimed, and what a practitioner would do with it.
What it actually is
Apex-G4 26B is not trained from scratch. It is the last link in a fine-tune chain that starts at Google. The card's declared lineage is google/gemma-4-26B-A4B → google/gemma-4-26B-A4B-it → TrevorJS/gemma-4-26B-A4B-it-uncensored → the BugTraceAI DPO fine-tune, with the declared base being TrevorJS/gemma-4-26B-A4B-it-uncensored.
The Google model underneath is a Mixture-of-Experts: 25.2B total parameters, 3.8B active (8 of 128 experts per token), 30 layers, a 256K context window, and hybrid local-sliding + global attention. The "26B" in the name rounds the total; only ~3.8B parameters are active per token, so its inference cost is closer to a 4B dense model than to a 26B one.
The intermediate step is real and worth naming. TrevorJS/gemma-4-26B-A4B-it-uncensored is an abliterated build of Gemma 4 — refusals removed via a stated "norm-preserving biprojected abliteration" plus "Expert-Granular Abliteration." The refusal removal that defines this model happened one layer upstream, before BugTraceAI's own training.
It is distributed as GGUF: an f16 "Master" file at 50.5 GB and a Q4_K_M "TurboQuant" at 16.7 GB, so it runs under llama.cpp / Ollama rather than transformers-natively.
Two structural changes are asserted on the card with no supporting detail: that the Gemma vision tower was stripped to make the model text-only, and that an "Opus-style reasoning engine" was injected to force step-by-step reasoning inside <thinking> blocks. Both are author claims with no technical documentation.
The family
Apex-G4 26B is not a standalone release. It is the largest member of the BugTraceAI family — four fine-tunes from the same anonymous account, each on a different open base, shipped across Q4, Q6, and F16 quant tiers.
| Model | Publisher-stated base | Hugging Face |
|---|---|---|
| CORE-Fast | Qwen2.5-Coder 7B | BugTraceAI-CORE-Fast |
| CORE-Pro | Mistral-Nemo 12B | BugTraceAI-CORE-Pro |
| Apex-G4 (this post) | publisher-described 26B Gemma MoE | Q4 · Master f16 |
| CORE-Ultra | Qwen3.6 27B | Q4 · Q6 |
All four are marketed for the same work — exploit chains, vulnerability research, and WAF/EDR/AV evasion — and all four carry the same evidentiary gap. The base attributions above are the publisher's own, no paper documents any member of the family, and the capability claims are self-reported. The scrutiny this post applies to Apex-G4 applies to the whole line: a stated base, a stated purpose, and no independent evaluation of any of them.
Claims vs. what is verified
This is the section that matters for this model, because the split is stark.
Traceable to a primary source:
- The provenance chain and declared base model — from the BugTraceAI Q4 card metadata.
- The base Gemma 4 architecture and its published benchmarks — from Google's model card.
- That the intermediate is an abliterated Gemma 4 — from the TrevorJS card.
- The GGUF formats and file sizes — from the Master card.
- The Apache-2.0 tag, consistent with Google releasing Gemma 4 under Apache 2.0.
Author claims, not independently verified:
- DPO training on a private "Super Dataset" of bug-bounty reports, malware methodologies, and WAF-evasion techniques. No data card, size, or provenance is published.
- The "injected reasoning engine" and "stripped vision tower."
- The single performance figure: 100% offensive compliance / 0% refusal on CyberSecEval / PurpleLlama's MITRE ATT&CK set. It is self-reported, and it measures willingness to answer — not correctness or exploit quality.
- The creator. BugTraceAI is an anonymous Hugging Face account with no disclosed team or affiliation. Mirror repos under other accounts are re-uploads, not validators, and the secondary write-ups that exist (awesomeagents.ai, a Medium post) appear to re-report or originate from the author rather than to test the model independently.
Benchmarks
No independent, third-party benchmark of the BugTraceAI fine-tune exists. Every strong number associated with this model belongs to the base Gemma 4, and the only number for the fine-tune itself is a self-reported compliance rate. The distinction is the whole point of the table.
| Benchmark | Score | Measured on |
|---|---|---|
| MMLU Pro | 82.6% | BASE — google/gemma-4-26B-A4B-it |
| GPQA Diamond | 82.3% | BASE — google/gemma-4-26B-A4B-it |
| AIME 2026 (no tools) | 88.3% | BASE — google/gemma-4-26B-A4B-it |
| LiveCodeBench v6 | 77.1% | BASE — google/gemma-4-26B-A4B-it |
| Codeforces (ELO) | 1718 | BASE — google/gemma-4-26B-A4B-it |
| CyberSecEval / MITRE ATT&CK — compliance / refusal | 100% / 0% | FINE-TUNE, self-reported — BugTraceAI card |
Read the last row carefully. A 100% compliance / 0% refusal figure tells you the model answers every MITRE ATT&CK prompt. It says nothing about whether the answers are correct, novel, or functional. There is no accuracy, pass@k, or functional-validity benchmark for this model — and none for the 4-bit quant specifically.
For a practitioner
What it gives you over stock Gemma 4. One thing, demonstrably: it will not refuse technically grounded offensive requests. The author positions it for "designing complex exploit chains and multi-stage payloads," "WAF/EDR/AV Evasion," and "Malware Analysis & Development". What is actually shown is compliance, delivered by the upstream abliteration. Its raw reasoning and coding ability is Gemma 4's — the base benchmarks above are the best available proxy, and they say nothing about the fine-tune's added value on real exploit work.
How to run it. It is GGUF, so llama.cpp / Ollama; the card ships an Ollama-style Modelfile. The f16 is 50.5 GB and the Q4_K_M is 16.7 GB. It inherits Gemma 4's 256K context, but the card recommends capping to 8192–16384 tokens and markets the Q4 as fitting a 12 GB card (e.g. an RTX 3060) via CPU offload. Recommended sampling is temperature 0.1, top_p 0.9, repeat penalty 1.1. The prompt convention is to open with a <thinking> block; expect verbose chain-of-thought before any payload. Because it is MoE with ~3.8B active parameters, throughput is closer to a 4B model than a dense 26B.
The real caveats.
- The only number is a compliance rate. You are the validation layer — the model does not tell you the exploit compiles, the payload evades, or the technique is current.
- Guardrails are deliberately gone. Refusal removal is the design goal, applied by upstream abliteration; outputs are unvetted. Apache-2.0 does not sanction unlawful use, and the uncensored intent shifts legal and ethical responsibility to the operator.
- Conversion quality is unmeasured. No perplexity or regression numbers exist for the fine-tune, so how much the abliteration + DPO + f16 GGUF pipeline moved base quality is unknown.
Honest limits
What is genuinely not known about this model:
- Whether the fine-tune improves on Gemma 4 at any real offensive task — no independent evaluation exists.
- What it was trained on — the "Super Dataset" has no published data card, size, or provenance.
- Whether the "injected reasoning engine" and "stripped vision tower" are real changes or marketing.
- Who made it — the author is anonymous with no track record.
- How much the quantization and abliteration degraded the base — unbenchmarked.
That is the honest shape of it: a competent open MoE with its refusals removed, one self-reported willingness metric, and no capability evidence of its own. For offensive research it may reduce friction. It does not, on the record, make anything more correct.
BugTraceAI Apex-G4 26B sits in the AdversariaLLM catalog as a measured entry — provenance, claims, and that one self-reported number, kept separate from Gemma 4's own: /catalog/bugtraceai-apex-g4-26b. For how it fits the wider field of open security-research models in 2026, see the survey: /blog/open-security-research-llms-2026.