All resources
Topic

eval-harness

3 resources across 1 kinds

Benchmarks

  1. Common evaluation harness/plumbing (from NIST CAISI) for running CVE-Bench, Cybench, and related environments rather than new training content.

    Open ↗
  2. MIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.

    Open ↗
  3. jev-sec-benchcloud costhosted

    MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.

    Open ↗