Benchmark · eval-harnesscloud costhosted
jev-sec-bench
MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.
Use responsibly
Deploys real cloud resources in your own account — this can cost money. Tear it down after use.
More benchmarks
- Jev Prompt-Injection Benchmark (JurijZ)eval-harnessMIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.
- NIST CAISI Cyber Evalseval-harnessCommon evaluation harness/plumbing (from NIST CAISI) for running CVE-Bench, Cybench, and related environments rather than new training content.
- BountyBenchbug-bountyBug-bounty evaluation of 40 historical bounties across 25 real systems with separate Detect, Exploit, and Patch tasks and executable graders.
- CVE-Benchweb-cveForty containerized critical real-world web CVEs with deterministic/executable graders for testing verified completion.
- CWE-Bench-Javawhite-box-javaJava white-box discovery/exploitation/patching benchmark from the IRIS-SAST project.
- Cybenchctf-long-horizonCTF benchmark for long-horizon planning and recovery.