All resources
Topic

safety-judge

1 resources across 1 kinds

Frameworks & agents

  1. A Python library providing standardized, tested implementations of 15+ LLM safety judges (StrongREJECT, LlamaGuard, WildGuard, HarmBench, and others) behind a unified API. Returns a normalized 0-1 harm score (p_harmful) for LLM conversations, supports both local fine-tuned and remote foundation-model judges, and warns when a setup diverges from the original implementation to preserve reproducibility.

    Open ↗