LLM4Decompile 6.7B

binary/asm → C

A specialized decompilation model (binary/assembly → C, refining Ghidra output) over DeepSeek-Coder-6.7B. A task model, not a general chat model.

Context4,096 (unconfirmed)
Curationcurated
AvailabilityComing soon
Price / Mtoknot yet priced
BaseDeepSeek-Coder-6.7B
QuantizationQ4_K_M
LicenseMIT

Source: LLM4Binary/llm4decompile-6.7b-v2 on HuggingFace

Good for
  • +decompilation
  • +binary → C
  • +Ghidra refinement
Model cardfrom the published weights — parameters, architecture, training, intended use
Parameters6.7B
ArchitectureDeepSeek-Coder 6.7B (LlamaForCausalLM)
Modalitytext
Training

Trained for assembly→C decompilation (O0–O3) over DeepSeek-Coder-6.7B (LLM4Binary).

Intended use
  • ›Decompile / refine x86 binary code into human-readable C source for
  • ›Refine Ghidra headless pseudo-code into C via the card's documented
  • ›Handle individual functions compiled at GCC optimization levels O0,
  • ›Programmatic use through transformers AutoModelForCausalLM (bfloat16
Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

LLM4Decompile aims to decompile x86 assembly instructions into C. This repository is the V2 6.7B model (the "LLM4Decompile-Ref-6.7B" refinement variant), which the card describes as "trained with a larger dataset (2B tokens) and a maximum token length of 4,096, with remarkable performance (up to 100% improvement) compared to the previous model." In the card's documented workflow it is used as a refinement (Ref) model, not a raw disassembler: a binary is first decompiled with Ghidra 11 headless, the target function's Ghidra pseudo-code is extracted and wrapped in a fixed prompt, then passed to the model, which emits refined C source. It is a single-purpose completion model rather than a chat model — the input/output format is fixed and it decompiles one function at a time. The card reports re-executability-rate and edit-similarity metrics across GCC optimization levels O0-O3, where the Ref-6.7B model beats the Ghidra baseline and Ghidra+GPT-4o on the reported numbers.

Training data

Card states the V2 series were "trained with a larger dataset (2B tokens) and a maximum token length of 4,096." No specific dataset name, source corpus, or license of the training binaries is given in the card. Base/architecture is deepseek-ai/deepseek-coder-6.7b-base (from config.json). It is a causal-LM decompiler; in the documented pipeline this V2 model is the "Ref" (refinement) variant that takes Ghidra headless pseudo-code as input and refines it into C.

Intended use
  • +Decompile / refine x86 binary code into human-readable C source for reverse-engineering and binary-analysis research.
  • +Refine Ghidra headless pseudo-code into C via the card's documented V2 pipeline (Ghidra 11.0.3 + Java-SDK-17 preprocessing, then the model as the refinement step).
  • +Handle individual functions compiled at GCC optimization levels O0, O1, O2, O3.
  • +Programmatic use through transformers AutoModelForCausalLM (bfloat16 on CUDA) with the fixed decompile prompt, as shown in the card.
Limitations
  • −Not a chat model: it uses a fixed completion prompt ('# This is the assembly code:' ... '# What is the source code?'); no system/chat template is provided.
  • −Max token length is 4,096 (card); the card notes max_new_tokens should stay below that range, so long functions can be truncated.
  • −Decompiles one function at a time — the card explicitly notes 'we only decompile one function, where the original file may contain multiple functions.'
  • −Requires the Ghidra headless preprocessing pipeline (Ghidra 11.0.3 + Java-SDK-17) to generate the pseudo-code that this V2 model refines.
  • −Scope is x86 assembly -> C per the card.
  • −Correctness is not guaranteed: the best reported average re-executability is 0.5274 (Ref-6.7B), i.e. a large fraction of outputs are not re-executable, and edit similarity to ground-truth source is low (~0.14 avg).
  • −The card contains no explicit limitations, safety, bias, or out-of-scope section; the constraints above are drawn from the card's usage notes and reported metrics.
Running the weights yourself

The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_path = 'LLM4Binary/llm4decompile-6.7b-v2' # V2 Model
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.bfloat16).cuda()

with open(fileName +'_' + OPT[0] +'.pseudo','r') as f:#optimization level O0
    asm_func = f.read()
inputs = tokenizer(asm_func, return_tensors="pt").to(model.device)
with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=2048)### max length to 4096, max new tokens should be below the range
c_func_decompile = tokenizer.decode(outputs[0][len(inputs[0]):-1])

with open(fileName +'_' + OPT[0] +'.pseudo','r') as f:#original file
    func = f.read()

print(f'pseudo function:\n{func}')# Note we only decompile one function, where the original file may contain multiple functions
print(f'refined function:\n{c_func_decompile}')
Benchmarks — as reported on the card

Reported by the model’s authors, not our own testing — our scores are in the table above.

Re-executability Rate, AVG O0-O3 (LLM4Decompile-Ref-6.7B = this repo)0.5274Vendor/author self-reported. Per level: O0 0.7439 / O1 0.4695 / O2 0.4756 / O3 0.4207. This repo is the '+LLM4Decompile-Ref-6.7B' row, which refines Ghidra pseudo-code. No benchmark dataset is named in the card.
Edit Similarity, AVG O0-O3 (LLM4Decompile-Ref-6.7B = this repo)0.1382Vendor self-reported. Per level: O0 0.1559 / O1 0.1353 / O2 0.1342 / O3 0.1273.
Re-executability Rate, AVG (Ghidra baseline)0.2012Vendor-reported comparator (Ghidra alone). Per level: O0 0.3476 / O1 0.1646 / O2 0.1524 / O3 0.1402.
Re-executability Rate, AVG (Ghidra + GPT-4o)0.3522Vendor-reported comparator. Per level: O0 0.4695 / O1 0.3415 / O2 0.2866 / O3 0.3110.
Re-executability Rate, AVG (LLM4Decompile-End-6.7B)0.4537Vendor-reported. The 'End' variant decompiles assembly directly (a different model, not this repo). Per level: O0 0.6805 / O1 0.3951 / O2 0.3671 / O3 0.3720.
Re-executability Rate, AVG (LLM4Decompile-Ref-33B)0.5091Vendor-reported context: the larger 33B Ref model scores slightly below this 6.7B Ref model on the AVG re-executability metric.
Scores

No measurements published for this version yet.

Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

Versionsscores attach to a version; v2 does not inherit v1's numbers
v12026-08-07—current

Start with LLM4Decompile 6.7B

A confirmed email account includes 30 free messages a month.

Sign in to start
LLM4Decompile 6.7B — Decompilation · AdversariaLLM