evaluation
8 resources across 3 kinds
Frameworks & agents
- Open ↗
A Python library providing standardized, tested implementations of 15+ LLM safety judges (StrongREJECT, LlamaGuard, WildGuard, HarmBench, and others) behind a unified API. Returns a normalized 0-1 harm score (p_harmful) for LLM conversations, supports both local fine-tuned and remote foundation-model judges, and warns when a setup diverges from the original implementation to preserve reproducibility.
Benchmarks
- Open ↗
An OpenSSF benchmark containing code and metadata for over 200 real-life CVEs, with tooling to evaluate Static Analysis Security Testing (SAST) solutions by measuring vulnerability detection rates and false positives against vulnerable/patched open-source code.
- Open ↗Jev Prompt-Injection Benchmark (JurijZ)cloud costpassive
MIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.
- Open ↗jev-sec-benchcloud costhosted
MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.
References
- Open ↗In-The-Wild Jailbreak Prompts on LLMsdual-usepassive
Dataset of 15,140 in-the-wild prompts collected from Reddit, Discord, websites and open-source datasets, 1,405 of them jailbreaks, plus a 390-question forbidden-scenario set for measuring jailbreak effectiveness.
- Open ↗Jailbreaking Frontier Modelsactivecloud costdual-usehigh-risklicence
Dataset of harmful-behaviour prompts (drug, chemical, biological, radiological, nuclear, explosive) and a reference PRBO reward function for training jailbreaking agents; the RL training loop is not included.
- Open ↗[un]prompted 2026 slide archivedual-uselicence
Community GitHub archive of 49 slide decks from [un]prompted 2026, the AI Security Practitioner Conference (March 3-4, San Francisco), spanning AI governance, agent security, offensive AI and agent evaluation; no licence stated.