Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪
-
Updated
Sep 11, 2026 - Python
Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪
Simple Prompt Injection Kit for Evaluation and Exploitation
[NDSS'25 Best Technical Poster] A collection of automated evaluators for assessing jailbreak attempts.
A working POC of a GPT-5 jailbreak via PROMISQROUTE (Prompt-based Router Open-Mode Manipulation) with a barebones C2 server & agent generation demo.
First-of-its-kind AI benchmark for evaluating the protection capabilities of large language model (LLM) guard systems (guardrails and safeguards)
LMAP (large language model mapper) is like NMAP for LLM, is an LLM Vulnerability Scanner and Zero-day Vulnerability Fuzzer.
Implementation of paper 'Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing'
[ICML 2025] Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
Mechanism-grounded taxonomy of 40 LLM jailbreak patterns across 10 categories. 8,000-trial bootstrap evaluation for the June 2026 frontier (Claude Opus 4-8, GPT-5.5, Gemini 3.5, DeepSeek V4). Every citation direct-WebFetch verified; refuted claims documented.
RetardBench is an open, no-censorship benchmark that ranks large language models purely on how retarded they are.
Benchmark LLM jailbreak resilience across providers with standardized tests, adversarial mode, rich analytics, and a clean Web UI.
Jailbreak Evaluation Framework -- 2025 Graduate Design for HFUT
Debugged version for Tree of Attacks: Jailbreaking Black-Box LLMs Automatically paper and added GPU optimization.
Chain-of-thought hijacking via template token injection for LLM censorship bypass (GPT-OSS)
Bypass all Google Gemini safety filters for uncensored text and image generation with a persistent jailbreak method.
LLM Jailbreaking via Prompt Rewriting
The Self-Hosted AI Firewall & Gateway. Drop-in guardrails for LLMs running entirely on CPU. Blocks jailbreaks, enforces policies, and ensures compliance in real-time
PESU I/O The Hacker's Gauntlet 24-hours CTF
Detect and prevent large language model jailbreaks using hidden state causal monitoring to enhance security in AI applications.
Build a runtime for Claude Code with persistence, observability, and multi-process coordination for long-running AI agents
To associate your repository with the llm-jailbreaks topic, visit your repo's landing page and select "manage topics."