AI-powered CTF assistant CLI featuring progressive RAG hints, a LangGraph agent, and deep Kali Linux integration via MCP.
CTF-GPT is designed to act as your pair-hacking operator. In Command Mode, it provides accurate, specific, ready-to-execute commands grounded in real CTF writeups to help you advance. In Agent/Auto Modes, it can execute tool workflows on your Kali VM (via MCP), then uses the outputs to update an evidence blackboard and produce grounded, context-aware next steps.
- Accurate Command Generation: Ask for commands and get accurate, specific methodology to keep moving. The response dynamically scales from concise one-liners to step-by-step commands based on query complexity.
- Agent Mode with Kali MCP: The assistant can run
nmap,gobuster,sqlmap, and more directly on your Kali machine, analyzing the output to update its evidence blackboard. - Pentesting Task Tree (PTT): Automatically tracks progress across structured phases (
recon→enumeration→exploitation→post-exploitation). - Tool Output Parsers: Condenses verbose tool outputs to extract structured findings, saving up to 80% in LLM token usage.
- Web Intelligence Fallback: Seamlessly pulls data from ExploitDB, CVE/NVD, and HackTricks when the local RAG knowledge base is empty.
- Session Memory: Learns from successful engagements by saving attack patterns, automatically applying this knowledge to similar challenges in the future.
- RAG Pipeline: Grounded in thousands of CTF writeups ingested from CTFtime, GitHub, and HackTricks.
- Auto-Reports: Generates a detailed Markdown report of your agent session, summarizing findings, dead ends, and executed commands.
- Language: Python 3.11+
- CLI Framework: Typer & Rich
- AI Orchestration: LangChain & LangGraph
- Vector Store: ChromaDB
- Embeddings: Sentence Transformers (
all-MiniLM-L6-v2) - System Integration: Model Context Protocol (MCP)
- Python 3.11 or higher
- An API Key for a cloud provider:
- Groq (
GROQ_API_KEY) OR - DeepSeek (
DEEPSEEK_API_KEY)
- Groq (
- Optional: Ollama installed (if running entirely local)
- Optional: A Kali Linux environment (for Agent Mode)
If you are running this directly on your Kali Linux machine (recommended for Agent mode):
git clone https://github.com/XploitMonk0x01/ctfgpt
cd ctfgptIt is highly recommended to use a virtual environment:
python3 -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
pip install -e .Set your API keys as environment variables. For example, using Groq:
export GROQ_API_KEY="<your_groq_api_key>"
# or for DeepSeek:
export DEEPSEEK_API_KEY="<your_deepseek_api_key>"By default, CTF-GPT is configured for Groq. You can verify your configuration with:
ctfgpt config --showIf you want to use DeepSeek instead:
ctfgpt config --set cloud.provider --value deepseekIf you want to switch to a different model on your provider (for example, Qwen on Groq):
ctfgpt config --set cloud.model --value "qwen/qwen3-32b"Before you can ask for commands, CTF-GPT needs data. Ingest writeups from various sources:
# Ingest the latest 100 writeups from CTFtime
ctfgpt ingest --source ctftime --limit 100
# Ingest from GitHub repositories
ctfgpt ingest --source github --repos "w181496/Web-CTF-Cheatsheet" --limit 50Verify your ingestion status:
ctfgpt statusAsk the assistant for accurate commands regarding a specific challenge:
# Get specific commands to run
ctfgpt ask "I have a PNG file but running strings shows a PK header at the top"
# In the interactive loop, you can follow up:
# Continue with a follow-up query or exit?
> "How do I extract it?"To use Agent Mode, CTF-GPT needs to communicate with Kali Linux via the Model Context Protocol (MCP).
If you cloned this repo directly onto your Kali Linux machine, setup is simple:
# Install the official mcp-kali-server package (assuming it's in the apt repo)
sudo apt update
sudo apt install mcp-kali-server
# Start the server
kali-server-mcp --ip 127.0.0.1 --port 5000ctfgpt config --set mcp.enabled --value true
ctfgpt config --set mcp.host --value localhostNote: If you are running CTF-GPT on Windows and Kali in a VM, you will need to setup an SSH tunnel:
ssh -L 5000:localhost:5000 kali@<KALI_IP>
Once connected, ask the agent to investigate:
ctfgpt ask "The target is 10.10.11.230. Start recon." --agentThe agent will autonomously plan its approach, but will prompt you for approval before executing each command on Kali.
Once you approve, it executes the tool (like nmap), reads the output, updates its evidence blackboard, and repeats the cycle until it has enough evidence to provide you with grounded commands.
After the session, view the generated report:
ctfgpt reportctfgpt/
├── ctfgpt/
│ ├── agent.py # LangGraph StateGraph definition
│ ├── blackboard.py # Pheromone-weighted evidence tracker
│ ├── classifier.py # Auto-detects CTF categories (web, pwn, etc.)
│ ├── cli.py # Typer CLI entrypoints
│ ├── config.py # Configuration & LLM instantiation
│ ├── mcp_client.py # Thin HTTP wrapper for Kali MCP server
│ ├── memory.py # Session memory for learning attack patterns
│ ├── parsers.py # Structured tool output extraction
│ ├── rag.py # LangChain LCEL Retrieval logic
│ ├── report.py # Jinja2 markdown report generator
│ ├── task_tree.py # Pentesting Task Tree (PTT) tracker
│ ├── web_intel.py # ExploitDB & CVE fallback intelligence
│ └── utils/
│ ├── history.py # Session history management
│ ├── rich_output.py# Terminal styling and UI
│ └── safety.py # Command validation and scope enforcement
├── ingestion/
│ ├── chunker.py # Semantic chunking and metadata tagging
│ ├── embedder.py # ChromaDB insertion logic
│ ├── loader_github.py # GitHub writeup loader
│ ├── loader_pdf.py # PDF writeup loader
│ ├── scraper_ctftime.py# CTFtime web scraper
│ └── scraper_hacktricks.py # HackTricks wiki loader
├── config.yaml # Default configuration
└── pyproject.toml # Project metadata and dependencies
graph TD
USER["👤 User Terminal"] --> CLI["CLI Layer<br/>Typer + Rich"]
CLI -->|"hint mode"| CLASSIFY["Category Classifier"]
CLI -->|"agent mode"| CLASSIFY
CLASSIFY --> RAG["RAG Chain<br/>LangChain LCEL"]
RAG --> CHROMA["ChromaDB<br/>6 Collections"]
RAG --> LLM["LLM Layer<br/>Groq / DeepSeek"]
CLASSIFY -->|"agent mode"| AGENT["LangGraph Agent<br/>StateGraph 5 nodes"]
AGENT --> BB["Session Blackboard<br/>JSON + Pheromones"]
AGENT --> MCP["MCP Client<br/>~40 lines HTTP"]
AGENT --> RAG
MCP -->|"HTTP"| KALI["mcp-kali-server<br/>Kali Package"]
KALI --> TOOLS["nmap · gobuster · hydra<br/>john · sqlmap · nikto<br/>metasploit · shell"]
BB -->|"findings"| AGENT
AGENT --> REPORT["Report Generator<br/>Jinja2 → Markdown"]
INGEST["Ingestion Pipeline"] --> CHROMA
| Command | Description |
|---|---|
ctfgpt ask "query" |
Get accurate commands and attack methodology in an interactive loop. Use --agent to run Kali tools. |
ctfgpt ask "query" --agent |
Agent Mode (executes whitelisted tools via Kali MCP). Supports --dry-run, --scope, --max-iter. |
ctfgpt solve <target> |
Run a category-aware predefined playbook (with per-step approval). Supports --category, --file, --max-steps, --dry-run, --scope. |
ctfgpt plan <target> |
Generate and execute an adaptive LLM plan (approval-based execution). Supports --category, --file, --max-steps. |
ctfgpt auto <target> |
Fully autonomous multi-agent run (Router + Recon/Exploit/PrivEsc). |
ctfgpt ingest |
Scrape + chunk + embed CTF writeups into ChromaDB. Supports --source, --limit, and GitHub/PDF options. |
ctfgpt status |
Check ChromaDB, LLM connectivity, and MCP server health. |
ctfgpt config |
View config (--show) or set values (--set key --value val). |
ctfgpt history |
Show session history (best-effort; Phase 2 placeholder). |
ctfgpt report |
Generate/view a session report. Supports --session, --list, --open. |
ctfgpt tools |
List MCP tools + connection status. |
ctfgpt solve is a smarter, more opinionated alternative to --agent. Instead of letting the LLM freely decide what to run, it executes a pre-defined playbook of the best tools for each CTF category, in the optimal order.
| Category | Tools Executed (in order) |
|---|---|
web |
curl headers → nikto → gobuster → robots.txt → source hints |
forensics |
file → strings → xxd → binwalk → exiftool → steghide |
pwn |
file → checksec → strings → nm → objdump → ltrace |
reversing |
file → strings → readelf → nm → objdump → anti-debug check |
crypto |
base64 decode → hex decode → ROT13 → Caesar brute → hashid |
osint |
whois → nslookup → curl headers → gobuster dns → openssl cert |
# Web target — auto-detects category, runs web playbook
ctfgpt solve http://10.10.11.230
# Force forensics category on a file
ctfgpt solve /home/kali/ctf/mystery.png --category forensics
# Crypto challenge — brute-force common encodings
ctfgpt solve "KHOOR ZRUOG" --category crypto
# Preview all planned steps without executing
ctfgpt solve 10.10.11.230 --dry-run
# Run only the first 3 steps
ctfgpt solve 10.10.11.230 --max-steps 3ctfgpt plan is the most intelligent attack mode. Unlike solve (static playbook) or --agent (open-ended), it asks the LLM to generate a concrete attack plan first, shows it to you, then executes step by step with adaptive re-planning after each tool output.
- Plan Generation — LLM creates a numbered attack plan based on your target, category, and RAG context
- User Approval — Full plan displayed in a table for review before execution
- Adaptive Execution — After each step, the LLM evaluates the output and can:
CONTINUE— proceed to the next stepINSERT— add an urgent new step (e.g., update/etc/hostsafter a redirect)REPLAN— revise the entire remaining plan based on new evidenceDONE— stop early if the flag is found
- Final Summary — RAG-grounded solution summary from all evidence
# Web target with adaptive planning
ctfgpt plan 10.10.11.230 --category web
# Let it auto-detect category from description
ctfgpt plan "WordPress site at smol.thm with vulnerable plugins"
# Interactive mode (prompts for target)
ctfgpt planctfgpt auto unleashes the Multi-Agent System. Instead of a single agent doing everything, the system spins up specialized sub-agents:
- Router: Analyzes the evidence and delegates tasks to the best sub-agent
- Recon Agent: Focused entirely on safe enumeration (nmap, gobuster, etc.)
- Exploit Agent: Focused on gaining initial access (sqlmap, msfconsole, etc.)
- PrivEsc Agent: Focused on local privilege escalation to root (linpeas, sudo, etc.)
Each agent runs its own ReAct loop with a strictly enforced whitelist of allowed tools. They share state through a unified Blackboard.
ctfgpt auto 10.10.11.230Error: ctfgpt status shows LLM disconnected.
Check:
GROQ_API_KEY(Groq) orDEEPSEEK_API_KEY(DeepSeek) is setctfgpt config --showpoints to the correct provider/model
Error: Agent mode fails or ctfgpt status shows MCP disconnected.
Check:
ctfgpt config --set mcp.enabled --value truectfgpt config --set mcp.host --value <host>kali-server-mcpis reachable onmcp.host:mcp.port(default port: 5000)- If running across a VM/host, ensure your port-forward / SSH tunnel is active
Error: SQLite/ChromaDB throws operational errors.
Solution: The database is stored under ~/.ctfgpt/db. If it becomes corrupted, delete it and re-run ctfgpt ingest.