Multi-hop cross-prompt injection benchmark for multi-agent AI systems. 250 attack cases, 8 taxonomy categories, 4 defenses evaluated. Watch: https://www.youtube.com/watch?v=fGOlMij4HPQ
-
Updated
Jun 16, 2026 - Python
Multi-hop cross-prompt injection benchmark for multi-agent AI systems. 250 attack cases, 8 taxonomy categories, 4 defenses evaluated. Watch: https://www.youtube.com/watch?v=fGOlMij4HPQ
A fail-closed runtime security layer for AI agents with prompt injection detection, tool-call protection, PII/secret sanitization, HITL controls, and output validation.
Hybrid AI Security Gateway — rule-based + LLM-assisted (Nemotron + Llama) prompt injection detection + MCP security (rules+policy+AI Assisted)
A practical repository showcasing guardrails for AI and LLM agents, including input/output validation, jailbreak protection, prompt security, and agent safety techniques to build secure and reliable AI systems.
Automated Prompt Injection Testing across 5 LLMs using Nvidia Garak
Open-source framework for evaluating the security of tool-using AI agents.
Defense-in-depth input safety for LLMs — perplexity gate + FAISS + ModernBERT + LoRA + Llama Guard 3, behind a deterministic policy gate. 99.88% accuracy, 99.47% jailbreak recall, calibrated confidence, ONNX-optimized. Live demo on HF Spaces.
Interactive LLM Attack & Defense Laboratory for demonstrating and analyzing LLM security threats.
Three-layer ensemble defence (anomaly + classifier + semantic) against prompt injection in RAG pipelines. F1=0.837, FPR=0.000.
Bitmask-based LLM security firewall with policy-driven logits filtering using GPU-accelerated Aho-Corasick pattern matching
To associate your repository with the llm-security-prompt-injection topic, visit your repo's landing page and select "manage topics."