eval-harness
3 resources across 1 kinds
Benchmarks
- Open ↗
Common evaluation harness/plumbing (from NIST CAISI) for running CVE-Bench, Cybench, and related environments rather than new training content.
- Open ↗Jev Prompt-Injection Benchmark (JurijZ)cloud costpassive
MIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.
- Open ↗jev-sec-benchcloud costhosted
MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.