All resources

Benchmarks

Held-out evaluations that measure real security capability.

12 shown
  1. BountyBenchbug-bounty

    Bug-bounty evaluation of 40 historical bounties across 25 real systems with separate Detect, Exploit, and Patch tasks and executable graders.

    Open ↗
  2. CVE-Benchweb-cve

    Forty containerized critical real-world web CVEs with deterministic/executable graders for testing verified completion.

    Open ↗
  3. CWE-Bench-Javawhite-box-java

    Java white-box discovery/exploitation/patching benchmark from the IRIS-SAST project.

    Open ↗
  4. Cybenchctf-long-horizon

    CTF benchmark for long-horizon planning and recovery.

    Open ↗
  5. CyberGymsource-analysis

    Large-scale benchmark for source analysis and PoC generation.

    Open ↗
  6. Common evaluation harness/plumbing (from NIST CAISI) for running CVE-Bench, Cybench, and related environments rather than new training content.

    Open ↗
  7. NYU CTF Benchctf-long-horizon

    NYU LLM CTF benchmark test set for long-horizon planning and recovery.

    Open ↗
  8. SEC-benchwhite-box

    Benchmark for white-box vulnerability discovery, exploitation, and patching.

    Open ↗
  9. XBOW-engineering validation benchmark set.

    Open ↗
  10. OpenSSF CVE Benchmarksource-analysis

    An OpenSSF benchmark containing code and metadata for over 200 real-life CVEs, with tooling to evaluate Static Analysis Security Testing (SAST) solutions by measuring vulnerability detection rates and false positives against vulnerable/patched open-source code.

    Open ↗
  11. Jev Prompt-Injection Benchmark (JurijZ)eval-harnesscloud costpassive

    MIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.

    Open ↗
  12. jev-sec-bencheval-harnesscloud costhosted

    MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.

    Open ↗