# BugTraceAI Apex-G4 26B: an abliterated Gemma 4 MoE with a compliance score, not a capability score

*D. Rose · 16 August 2026 · 6 min*

> A GGUF, DPO-fine-tuned, deliberately uncensored offensive-security build of Google's Gemma 4 26B MoE, shipped by an anonymous author. The one performance number attached to it measures whether the model answers — not whether its exploits work.

BugTraceAI Apex-G4 26B is a fine-tune of Google's Gemma 4 26B A4B, packaged as a GGUF and marketed for offensive research. It ships exactly one performance number, and that number measures whether the model answers a request — not whether what it produces compiles, evades, or works. Here is what is actually in the box, what is only claimed, and what a practitioner would do with it.

## What it actually is

Apex-G4 26B is not trained from scratch. It is the last link in a fine-tune chain that starts at Google. The card's declared lineage is `google/gemma-4-26B-A4B` → `google/gemma-4-26B-A4B-it` → `TrevorJS/gemma-4-26B-A4B-it-uncensored` → the BugTraceAI DPO fine-tune, with the declared base being [`TrevorJS/gemma-4-26B-A4B-it-uncensored`](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Q4/raw/main/README.md).

The Google model underneath is a Mixture-of-Experts: [25.2B total parameters, 3.8B active (8 of 128 experts per token), 30 layers, a 256K context window, and hybrid local-sliding + global attention](https://huggingface.co/google/gemma-4-26B-A4B-it). The "26B" in the name rounds the total; only ~3.8B parameters are active per token, so its inference cost is closer to a 4B dense model than to a 26B one.

The intermediate step is real and worth naming. [`TrevorJS/gemma-4-26B-A4B-it-uncensored`](https://huggingface.co/TrevorJS/gemma-4-26B-A4B-it-uncensored) is an [abliterated](/notes/what-abliteration-actually-does) build of Gemma 4 — refusals removed via a stated "norm-preserving biprojected abliteration" plus "Expert-Granular Abliteration." The refusal removal that defines this model happened one layer upstream, before BugTraceAI's own training.

It is distributed as GGUF: an [f16 "Master" file at 50.5 GB and a Q4_K_M "TurboQuant" at 16.7 GB](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16/raw/main/README.md), so it runs under llama.cpp / Ollama rather than transformers-natively.

Two structural changes are asserted on the card with no supporting detail: that the Gemma vision tower was stripped to make the model text-only, and that an ["Opus-style reasoning engine" was injected](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16/raw/main/README.md) to force step-by-step reasoning inside `<thinking>` blocks. Both are author claims with no technical documentation.

## The family

Apex-G4 26B is not a standalone release. It is the largest member of the **BugTraceAI** family — four fine-tunes from the same anonymous account, each on a different open base, shipped across Q4, Q6, and F16 quant tiers.

| Model | Publisher-stated base | Hugging Face |
|---|---|---|
| CORE-Fast | Qwen2.5-Coder 7B | [BugTraceAI-CORE-Fast](https://huggingface.co/BugTraceAI/BugTraceAI-CORE-Fast) |
| CORE-Pro | Mistral-Nemo 12B | [BugTraceAI-CORE-Pro](https://huggingface.co/BugTraceAI/BugTraceAI-CORE-Pro) |
| Apex-G4 (this post) | publisher-described 26B Gemma MoE | [Q4](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Q4) · [Master f16](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16) |
| CORE-Ultra | Qwen3.6 27B | [Q4](https://huggingface.co/BugTraceAI/BugTraceAI-CORE-Ultra-27B-Q4) · [Q6](https://huggingface.co/BugTraceAI/BugTraceAI-CORE-Ultra-27B-Q6) |

All four are marketed for the same work — exploit chains, vulnerability research, and WAF/EDR/AV evasion — and all four carry the same evidentiary gap. The base attributions above are the publisher's own, no paper documents any member of the family, and the capability claims are self-reported. The scrutiny this post applies to Apex-G4 applies to the whole line: a stated base, a stated purpose, and no independent evaluation of any of them.

## Claims vs. what is verified

This is the section that matters for this model, because the split is stark.

**Traceable to a primary source:**

- The provenance chain and declared base model — from the [BugTraceAI Q4 card metadata](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Q4/raw/main/README.md).
- The base Gemma 4 architecture and its published benchmarks — from [Google's model card](https://huggingface.co/google/gemma-4-26B-A4B-it).
- That the intermediate is an abliterated Gemma 4 — from the [TrevorJS card](https://huggingface.co/TrevorJS/gemma-4-26B-A4B-it-uncensored).
- The GGUF formats and file sizes — from the [Master card](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16/raw/main/README.md).
- The Apache-2.0 tag, consistent with Google releasing [Gemma 4 under Apache 2.0](https://opensource.googleblog.com/2026/03/gemma-4-expanding-the-gemmaverse-with-apache-20.html).

**Author claims, not independently verified:**

- DPO training on a private "Super Dataset" of bug-bounty reports, malware methodologies, and WAF-evasion techniques. No data card, size, or provenance is published.
- The "injected reasoning engine" and "stripped vision tower."
- The single performance figure: [100% offensive compliance / 0% refusal on CyberSecEval / PurpleLlama's MITRE ATT&CK set](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16/raw/main/README.md). It is self-reported, and it measures willingness to answer — not correctness or exploit quality.
- The creator. [BugTraceAI](https://huggingface.co/BugTraceAI) is an anonymous Hugging Face account with no disclosed team or affiliation. Mirror repos under other accounts are re-uploads, not validators, and the secondary write-ups that exist ([awesomeagents.ai](https://awesomeagents.ai/news/bugtraceai-apex-26b-local-red-team-model/), a [Medium post](https://albert-corzo.medium.com/we-built-an-offensive-security-ai-that-answered-every-mitre-att-ck-challenge-heres-what-happened-86c7ba1f7dab)) appear to re-report or originate from the author rather than to test the model independently.

## Benchmarks

No independent, third-party benchmark of the BugTraceAI fine-tune exists. Every strong number associated with this model belongs to the **base** Gemma 4, and the only number for the **fine-tune itself** is a self-reported compliance rate. The distinction is the whole point of the table.

| Benchmark | Score | Measured on |
|---|---|---|
| MMLU Pro | 82.6% | **BASE** — [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it) |
| GPQA Diamond | 82.3% | **BASE** — [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it) |
| AIME 2026 (no tools) | 88.3% | **BASE** — [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it) |
| LiveCodeBench v6 | 77.1% | **BASE** — [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it) |
| Codeforces (ELO) | 1718 | **BASE** — [google/gemma-4-26B-A4B-it](https://huggingface.co/google/gemma-4-26B-A4B-it) |
| CyberSecEval / MITRE ATT&CK — compliance / refusal | 100% / 0% | **FINE-TUNE**, self-reported — [BugTraceAI card](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16/raw/main/README.md) |

Read the last row carefully. A 100% compliance / 0% refusal figure tells you the model answers every MITRE ATT&CK prompt. It says nothing about whether the answers are correct, novel, or functional. There is no accuracy, pass@k, or functional-validity benchmark for this model — and none for the 4-bit quant specifically.

## For a practitioner

**What it gives you over stock Gemma 4.** One thing, demonstrably: it will not refuse technically grounded offensive requests. The author positions it for ["designing complex exploit chains and multi-stage payloads," "WAF/EDR/AV Evasion," and "Malware Analysis & Development"](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16/raw/main/README.md). What is actually shown is compliance, delivered by the upstream abliteration. Its raw reasoning and coding ability is Gemma 4's — the base benchmarks above are the best available proxy, and they say nothing about the fine-tune's added value on real exploit work.

**How to run it.** It is GGUF, so llama.cpp / Ollama; the card ships an Ollama-style Modelfile. The [f16 is 50.5 GB and the Q4_K_M is 16.7 GB](https://huggingface.co/BugTraceAI/BugTraceAI-Apex-G4-26B-Master-f16/raw/main/README.md). It inherits Gemma 4's 256K context, but the card recommends capping to 8192–16384 tokens and markets the Q4 as fitting a 12 GB card (e.g. an RTX 3060) via CPU offload. Recommended sampling is temperature 0.1, top_p 0.9, repeat penalty 1.1. The prompt convention is to open with a `<thinking>` block; expect verbose chain-of-thought before any payload. Because it is MoE with ~3.8B active parameters, throughput is closer to a 4B model than a dense 26B.

**The real caveats.**

- The only number is a compliance rate. You are the validation layer — the model does not tell you the exploit compiles, the payload evades, or the technique is current.
- Guardrails are deliberately gone. Refusal removal is the design goal, applied by upstream abliteration; outputs are unvetted. Apache-2.0 does not sanction unlawful use, and the uncensored intent shifts legal and ethical responsibility to the operator.
- Conversion quality is unmeasured. No perplexity or regression numbers exist for the fine-tune, so how much the abliteration + DPO + f16 GGUF pipeline moved base quality is unknown.

## Honest limits

What is genuinely not known about this model:

- Whether the fine-tune improves on Gemma 4 at any real offensive task — no independent evaluation exists.
- What it was trained on — the "Super Dataset" has no published data card, size, or provenance.
- Whether the "injected reasoning engine" and "stripped vision tower" are real changes or marketing.
- Who made it — the author is anonymous with no track record.
- How much the quantization and abliteration degraded the base — unbenchmarked.

That is the honest shape of it: a competent open MoE with its refusals removed, one self-reported willingness metric, and no capability evidence of its own. For offensive research it may reduce friction. It does not, on the record, make anything more correct.

---

BugTraceAI Apex-G4 26B sits in the AdversariaLLM catalog as a measured entry — provenance, claims, and that one self-reported number, kept separate from Gemma 4's own: [/catalog/bugtraceai-apex-g4-26b](/catalog/bugtraceai-apex-g4-26b). For how it fits the wider field of open security-research models in 2026, see the survey: [/blog/open-security-research-llms-2026](/blog/open-security-research-llms-2026).
