Benchmarks
Held-out evaluations that measure real security capability.
- Open ↗BountyBenchbug-bounty
Bug-bounty evaluation of 40 historical bounties across 25 real systems with separate Detect, Exploit, and Patch tasks and executable graders.
- Open ↗CWE-Bench-Javawhite-box-java
Java white-box discovery/exploitation/patching benchmark from the IRIS-SAST project.
- Open ↗NIST CAISI Cyber Evalseval-harness
Common evaluation harness/plumbing (from NIST CAISI) for running CVE-Bench, Cybench, and related environments rather than new training content.
- Open ↗NYU CTF Benchctf-long-horizon
NYU LLM CTF benchmark test set for long-horizon planning and recovery.
- Open ↗XBOW Validation Benchmarkssmoke-test
XBOW-engineering validation benchmark set.
- Open ↗OpenSSF CVE Benchmarksource-analysis
An OpenSSF benchmark containing code and metadata for over 200 real-life CVEs, with tooling to evaluate Static Analysis Security Testing (SAST) solutions by measuring vulnerability detection rates and false positives against vulnerable/patched open-source code.
- Open ↗
MIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.
- Open ↗
MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.