Curated archive of jailbreak and prompt-injection experiments for AI model robustness research, red-teaming, and defensive evaluation.
-
Updated
Aug 20, 2026
Curated archive of jailbreak and prompt-injection experiments for AI model robustness research, red-teaming, and defensive evaluation.
Enterprise LLM prompt security & red-teaming framework — automated jailbreak detection, prompt injection testing, policy enforcement, and comprehensive audit trails for production AI systems.
A structured NLP dataset for detecting prompt injection attacks, jailbreak attempts, and malicious instruction manipulation in Large Language Models (LLMs). Includes annotated threat categories, risk classifications, and validation-ready samples for AI safety training, security evaluation, and adversarial robustness research.
x402 settlement facilitator + EAS-compatible threat-intel attestation issuer on Base mainnet
Systematic red-teaming framework for adversarial prompt evaluation — jailbreak detection, injection classification, attack surface coverage metrics
PromptForge: LLM Red Teaming Toolkit is a lightweight toolkit for testing and hardening LLM applications.
To associate your repository with the adversarial-prompts topic, visit your repo's landing page and select "manage topics."