All resources
Topic

llm

21 resources across 5 kinds

Tools

  1. Client-side SillyTavern extension that rewrites outgoing Kimi/Moonshot chat requests to prefill an assistant reasoning_content turn, so the model continues thinking from user-seeded reasoning without a server patch.

    Open ↗
  2. jev-edgecloud cost

    Apache-2.0 fail-open gateway filter for OpenResty, APISIX, Kong, Envoy and JS hosts that scores inbound LLM requests for prompt injection, judging with a paid TypeSafe Jev key, a fine-tuned Laya or an OpenAI-compatible endpoint.

    Open ↗
  3. jevkit (Jev Agent Kit)cloud costhosted

    MIT-licensed CLI and MCP server giving agents eleven typed-decision tools on TypeSafe's hosted Jev model (route, triage, guard, grep, rank, compact), plus an advisory Claude Code PreToolUse hook; needs a TypeSafe API key.

    Open ↗

Frameworks & agents

  1. Pentest-Strategisttraining only

    A research framework that trains custom LLMs (Qwen 14B with reinforcement learning) on curated penetration-testing reasoning datasets to autonomously generate pentest strategies and step-by-step actions. Ships dataset-collection utilities, training code, RL experiments, and CTF-based evaluation, with an accompanying arXiv preprint.

    Open ↗
  2. An AI security co-pilot built on the open ZySec 7B model, trained across 30+ cybersecurity domains, providing threat analysis, playbook/document retrieval and standards reference for security professionals. Runs locally (CPU or GPU via vLLM) through a Streamlit UI and is OpenAI-API compatible.

    Open ↗
  3. Academic (TU Wien ipa-lab) open-source framework for building LLM-driven autonomous pentest agents in under ~50 lines, covering Linux privilege escalation, web-app and REST-API testing via SSH/local shell. Includes SQLite run logging and a web viewer for replaying agent runs; supports OpenAI and local models.

    Open ↗
  4. Open-source pentest assistant that pairs the WhiteRabbitNeo security LLM with the PentestGPT prompting methodology, using structured todo-list workflows and constrained (outlines) generation to keep the fully-open-source stack usable without proprietary APIs. Author notes parity with GPT-4 is still a work in progress.

    Open ↗
  5. Stop Sloppassive

    MIT-licensed skill file with phrase, structure and example references that teaches Claude or any LLM to strip recognisable AI-writing patterns from prose, with a five-dimension 1-10 scoring rubric; a writing aid, not a security tool.

    Open ↗

Labs & practice targets

  1. Vulnerable-LLM CTF challenges aligned to the OWASP Top 10 for LLM Applications, running a local Vicuna model (no cloud fees).

    Open ↗
  2. Sample LLM ReAct chatbot (Langchain) for learning prompt injection against Thought/Action/Observation agent loops.

    Open ↗

Benchmarks

  1. MIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.

    Open ↗
  2. jev-sec-benchcloud costhosted

    MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.

    Open ↗

References

  1. An empirical, practical guide to LLM hacking from security firm Forces Unseen, covering prompt injection, offensive and defensive techniques, and interactive playgrounds. Source for the handbook hosted at doublespeak.chat.

    Open ↗
  2. Collection of payloads for AI red teaming, organized around the OWASP AITG-APP (AI Testing Guide) categories. A reference set of adversarial prompt/test payloads for authorized LLM/AI application security testing rather than executable malware.

    Open ↗
  3. A curated index of resources on applying LLMs to automated penetration testing: 105+ academic papers grouped into systems/agents, benchmarks/cyber-ranges, empirical evaluations, surveys, and defense/ethics, plus links to code repositories (PentestGPT, VulnBot, Shannon and others) and evaluation benchmarks. Tied to the paper 'Hackers or Hallucinators?'.

    Open ↗
  4. MIT-licensed curated index of 225+ general-purpose AI tools (chat, coding, research, image, video, voice, agents, local model runners) with a per-tool writeup page; a general AI directory with no security-specific content.

    Open ↗
  5. MIT-licensed catalogue of open-source generative-AI SaaS templates (image, video, virtual try-on, writing, chatbots, voice) on a shared Next.js/Stripe/Google-OAuth stack, built to be forked and resold; no security content.

    Open ↗
  6. Dataset of 15,140 in-the-wild prompts collected from Reddit, Discord, websites and open-source datasets, 1,405 of them jailbreaks, plus a 390-question forbidden-scenario set for measuring jailbreak effectiveness.

    Open ↗
  7. Jailbreaking Frontier Modelsactivecloud costdual-usehigh-risklicence

    Dataset of harmful-behaviour prompts (drug, chemical, biological, radiological, nuclear, explosive) and a reference PRBO reward function for training jailbreaking agents; the RL training loop is not included.

    Open ↗
  8. README-only collection of copy-paste jailbreak prompts for DeepSeek R1, Grok 3, Gemini 2.0, ChatGPT (DAN), Claude 2 and Llama 2, plus a Gemini system-prompt leak prompt, mostly reposted from linked Reddit, blog and GitHub posts.

    Open ↗
  9. Vendor-organised collection of files, mostly Markdown, presenting what it says are the verbatim system prompts of major chatbots and coding agents (Anthropic, OpenAI, Google, xAI, Cursor, Kimi and others), with a dated additions table and an open invitation to PRs.

    Open ↗