Benchmark · ctf-long-horizon
Cybench
CTF benchmark for long-horizon planning and recovery.
Tagsctf-long-horizon
More benchmarks
- NYU CTF Benchctf-long-horizonNYU LLM CTF benchmark test set for long-horizon planning and recovery.
- BountyBenchbug-bountyBug-bounty evaluation of 40 historical bounties across 25 real systems with separate Detect, Exploit, and Patch tasks and executable graders.
- CVE-Benchweb-cveForty containerized critical real-world web CVEs with deterministic/executable graders for testing verified completion.
- CWE-Bench-Javawhite-box-javaJava white-box discovery/exploitation/patching benchmark from the IRIS-SAST project.
- CyberGymsource-analysisLarge-scale benchmark for source analysis and PoC generation.
- Jev Prompt-Injection Benchmark (JurijZ)eval-harnessMIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.