Live Demo: https://salmon-ground-0362fac00.7.azurestaticapps.net/
CodeSentinel is an advanced, autonomous AI agent mesh designed to identify, analyze, and remediate vulnerabilities and structural bugs within modern codebases. By leveraging a deterministic state machine (LangGraph) integrated with tiered Large Language Models (LLMs), CodeSentinel operates as a proactive, autonomous site reliability engineer (SRE) and security researcher. The system maps repositories, executes static analysis, generates AST-aware code patches, validates repairs against testing suites, runs targeted security re-verifications, and authors production-ready GitHub Pull Requests entirely autonomously.
Modern CI/CD pipelines are excellent at detecting issues via static application security testing (SAST), linting, and dependency scanning, but fundamentally fail at remediation. Developers are bombarded with alerts, leading to alert fatigue and growing technical debt. Critical fixes are delayed, and inconsistent patches arise when different engineers fix the same class of bug differently. Traditional AI coding assistants are reactive, requiring human-in-the-loop prompting and localized context gathering. There is a critical need for an asynchronous, autonomous pipeline capable of digesting global repository context, synthesizing architectural understanding, executing high-confidence code repairs, and verifying security fixes without manual intervention.
CodeSentinel bridges the gap between detection and remediation. When triggered, it performs a complete ingestion of the target repository. It utilizes a multi-agent architecture where highly specialized AI agents handle distinct phases of the remediation lifecycle:
- Contextual Mapping: Building an LLM-synthesized architectural mapping of the repository.
- Deep Investigation: Triaging anomalies detected by traditional static analyzers, utilizing Retrieval-Augmented Generation (RAG) against past successful fixes.
- Repair Planning: Formulating multi-file architectural repair strategies, with human-in-the-loop approval gates for high-risk operations (e.g., Auth, DB schemas).
- Code Generation: Applying precise, context-aware code patches.
- Validation: Verifying structural integrity post-repair via syntax checks and test suite executions.
- Security Verification: File-scoped, post-patch execution of security scanners (Bandit/Semgrep) to guarantee vulnerability closure.
- PR Authoring: Generating rich, conventional pull requests for human review.
CodeSentinel is built on a modern, decoupled architecture prioritizing extreme fault tolerance, real-time observability, and horizontal scalability.
-
Frontend Visualization Layer (React + Vite + Tailwind): A high-performance React application serving as the control plane. It consumes Server-Sent Events (SSE) to render real-time telemetry of the LangGraph execution state, providing operators with confidence scores, live diff rendering, and pipeline visualization.
- Admin Dashboard: An isolated, separate Vite/React application built for observability. It exposes macro-level system metrics, token rotation stats, queue depth, and job metrics. Accessing this dashboard requires the
X-Admin-Tokenheader.
- Admin Dashboard: An isolated, separate Vite/React application built for observability. It exposes macro-level system metrics, token rotation stats, queue depth, and job metrics. Accessing this dashboard requires the
-
API & Orchestration Layer (FastAPI + SQLite): A high-throughput asynchronous backend powered by FastAPI that acts as the Single Source of Truth for the entire system.
- Why the Monolithic Backend was Replaced: Originally, CodeSentinel packaged all language runtimes, build tools, and SAST scanners into a single 5.5GB monolithic backend container. This huge image made deploying to cost-effective serverless environments (like Azure Container Apps Free Tier) impossible due to strict 5-minute image pull timeouts.
- Why SQLite was Selected: SQLite provides robust, persistent job tracking and idempotent state management without the overhead, cost, or complexity of managing a dedicated database service (like PostgreSQL) for a lightweight orchestrator.
- Why Backend is the Single Source of Truth: The orchestrator tracks the state and broadcasts it via SSE. This ensures that the frontend never has to poll GitHub APIs directly, and the backend maintains authoritative control over ChromaDB, API keys, and job lifecycles.
-
Ephemeral Worker Layer (GitHub Actions):
- Why GitHub Actions was Chosen: We shifted the actual LangGraph execution and SAST scanning to GitHub Actions. This provides free, ephemeral, on-demand compute environments that come pre-installed with almost every language runtime and build tool imaginable.
- Agent Mesh (
backend/agents/): A swarm of stateless, specialized agents that execute within the GitHub Action. Each agent acts as a distinct node in the LangGraph, executing heavy scans and posting granular state updates back to the orchestrator via HTTP webhooks. - Repo Mapper: Builds a rich architectural map containing API endpoint inventories, database interaction mappings, and service boundary detection via LLM synthesis.
- Dependency Analyzer: Identifies outdated packages and CVEs (PyPI/npm/Maven/Go).
- Static Analysis: Scans for vulnerabilities across multiple categories including
security,quality(e.g., duplicate code, long methods), andperformance(e.g., SQLAlchemy N+1 detection). - Bug Investigator: Performs LLM RAG root-cause analysis querying historical fixes.
- Repair Planner: Formulates fixes & conditionally requests human approval for high-risk operations.
- Code Generator: Generates precise, AST-aware Search/Replace patches.
- Validator: Executes an automated build verification step followed by dynamic testing and security reverification.
- Security Verifier: Re-runs SAST on modified files to verify vulnerabilities are fixed.
- PR Author: Synthesizes diffs and interactions into a Pull Request for human review.
-
Intelligent Tooling & Routing (
backend/tools/): The proprietary LLM and operational infrastructure layer. Features a highly optimizedllm_routercapable of dynamic model tiering (fast/cheap vs. slow/reasoning), automatic JSON schema extraction, a highly resilientkey_dispatcherfor token budget load balancing, a unifiedanalysis_runnerfor executing core SAST tools (while additional scanners are invoked directly by the static analysis agent), aconfidence_calcengine for weighted pipeline scoring, and aknowledge_graphmodule that builds AST import dependency graphs and detects circular import cycles.
The data flow ensures context is preserved and strictly typed throughout the execution lifecycle.
- Ingestion & Triggering: A webhook or API call initiates the pipeline. The orchestrator initializes a strictly typed
PipelineStateobject containingrepo_url,retry_count, andconfidence_score. - Contextual Mapping: The
repo_mapperagent builds an LLM-synthesized architectural mapping of the repository, extracting boundaries and relationships. - Analysis Execution: The dependency_analyzer agent runs first, followed by the static_analysis agent, both injecting raw findings into the state using external tools (OSV, Bandit, Semgrep, ESLint, Pylint, Flake8, SonarQube (if available), Go Vet, Cargo Clippy) and real-time public registry checks (NPM, PyPI, Maven Central, Go Proxy, and Crates.io).
- LLM Triage & Context Retrieval: The
bug_investigatoragent queries the ChromaDB Vector Store for historical fixes of similar bugs, filtering out false positives using an LLM. - Planning & Approval: The
repair_plannerformulates a patch strategy. If modifications hit sensitive paths (e.g.,auth/,db/), it transitions the pipeline to anawaiting_approvalstate, broadcasting an SSE event and pausing the graph until a/api/approve/{task_id}webhook is received. - Execution Loop: The
code_generatorexecutes file I/O to apply patches. Thevalidatorinspects syntax and test outputs. If tests fail, the graph cycles back. - Security Verification Loop: Post-validation,
security_verifierexecutes targeted Semgrep/Bandit scans strictly on modified files. If the original vulnerability rule fires again, thesecurity_retry_contextis updated and the graph routes back tocode_generator. - Commit & PR: Upon successful validation and security clearance, the
pr_authorsynthesizes the diffs and interactions into a comprehensive Pull Request via the PyGithub SDK.
The LangGraph implementation uses a directed graph with tightly controlled cyclic capabilities for self-healing:
[ START ]
│
▼
( Repo Mapper ) ──► ( Dependency Analyzer ) ──► ( Static Analysis )
│
▼
( Bug Investigator )
│
▼
( Repair Planner ) ──► [ High Risk: Await Approval ]
│ │
[Low Risk] │ │ [Approved]
▼ ▼
┌◄───────────────────────── ( Code Generator ) ◄┐ ◄────────────┘
│ │ │
[ Validation Failed ] ▼ │
│ ( Validator ) │
└─────────────────────────────────────┤ │
│ │
[ Validation Passed ]
│ │
▼ │
( Security Verifier )┴─ [ Vulnerability Persists ]
│
[ Security Clean ]
│
▼
( PR Author )
│
▼
[ END ]
CodeSentinel optimizes cost and latency by routing tasks to appropriately sized models. Structural tasks (mapping, basic extraction) are routed to Tier 1 models (e.g., llama-3.1-8b-instant), while complex architectural reasoning and actual code synthesis are routed to Tier 2 models (e.g., llama-3.3-70b-versatile).
For maximum safety, repair_planner.py uses heuristic classify_risk checks. If high-risk changes are detected, execution pauses dynamically via asyncio.Event.wait(). The FastAPI backend streams an approval_required SSE event to the React frontend, displaying an Approval Modal. In the background worker (sse.py), asyncio.wait with FIRST_COMPLETED ensures telemetry and incoming webhooks don't cause thread deadlocks while the pipeline is streaming LangGraph state.
Beyond basic security checks, the static_analysis.py agent enforces high structural code quality across 21 specialized scanning modules (8 standard SAST tools and 13 custom modules). Findings are tagged with specific categories (security, quality, performance, functional).
- Quality & Dead Code: Enforces thresholds for long methods (50 lines) and duplicate code blocks, identifies circular dependencies via
knowledge_graph, detects built-in hardcoded secrets, and flags unused functions, classes, and files across Python, JS/TS, HTML, and CSS. - Performance & Leak Detection: Executes AST matching to detect N+1 query structures in SQLAlchemy, Prisma, Mongoose, Go, Rust, and Java. Identifies memory and resource leaks such as unclosed Python
open()calls, uncleaned JS/TS event listeners/intervals, Go unclosed resources, and Java unclosed streams.
Before running dynamic test suites, the validator.py agent executes a mandatory build verification step. It intelligently detects the project type (package.json, pom.xml, build.gradle, pyproject.toml, setup.py) and runs the associated build command (e.g., npm run build, mvn package, gradle assemble, python -m build). If the build fails, the pipeline immediately short-circuits back to the code_generator, saving valuable compute time that would otherwise be wasted on fundamentally broken syntax.
Rather than executing blind fixes, CodeSentinel ensures deterministic proof of remediation. The security_verifier.py isolates newly generated patches and runs semgrep and bandit strictly on modified files. If the patch fails to clear the initial rule ID, the system injects security_retry_context back into the graph, feeding the exact tool failure logs back to the code_generator for a secondary attempt.
CodeSentinel learns from its successes. Using confidence_calc.py, the validator.py and pr_author.py agents calculate a unified 4-part confidence score: (tests_ratio * 40.0) + (static_clean * 30.0) + (patches_ratio * 20.0) + (chroma_ratio * 10.0). Every successfully validated patch is hashed and embedded into a local validated_fixes ChromaDB collection. Future runs by the bug_investigator query this vector store via RAG to inject context from previously successful architectural repairs.
To bypass rate limits and token exhaustion, the system implements a proprietary load-balancer tracking daily token budgets across an array of API keys. Upon exhaustion or HTTP 429 errors, it seamlessly rolls over to the next key, eventually failing over to a highly restricted Emergency Key (configured via the GROQ_EMERGENCY_KEY environment variable).
To provide a massive leap in User Experience (UX), the FastAPI backend streams LangGraph state transitions directly to the React frontend via SSE. The pipeline execution is completely decoupled into a headless fastapi.BackgroundTasks runner generating a unique task_id. The event_generator drains background asyncio.Queue objects asynchronously based on this ID, allowing the UI to safely disconnect and reconnect without interrupting the robust CI/CD execution pipeline.
- Asynchronous Execution: Long-running AI agents are executed as background tasks within the asynchronous event loop. The
/api/analyzeendpoint immediately returns aTaskID, preventing HTTP timeouts. - RESTful + Streaming: Standard operations utilize REST, while active pipeline execution relies exclusively on SSE for real-time telemetry. This inherently supports strict corporate firewalls better than stateful WebSockets.
- Unblocking Mechanisms: The
/api/approve/{task_id}webhook interacts directly with the orchestration memory space to set thread events, elegantly waking sleeping graph nodes. - CI/CD Integration Pipeline: The primary deployment mechanism leverages GitHub Actions. A drop-in template (
codesentinel.yml) initiates the analysis via the/api/webhook/githubendpoint whenever a PR is opened or a push occurs. It validates the webhook payload via HMAC SHA-256 signature verification and responds by posting a comment on the PR containing a link to the live SSE telemetry stream.
- Stateless Orchestration & Ephemeral Compute: LangGraph nodes process state transformations on a
PipelineStateTypedDict entirely within an ephemeral GitHub Actions worker. This architecture inherently supports massive horizontal scaling because the heavy compute (cloning, building, scanning, LLM generation) is offloaded to GitHub's infrastructure, while the orchestrator remains an ultra-lightweight state machine and SSE broadcaster. - Aggressive Caching (
context_cache.py/response_cache.py): Identical LLM prompts and repository AST structures are cached locally. During repetitive debugging cycles, this prevents redundant network I/O to the LLM provider, drastically reducing pipeline execution time and API costs.
- Execution Environment: Code execution (e.g., running
npm run buildorpython -m buildduring the validation phase) currently runs directly on the host environment. Future updates plan to shift this execution into ephemeral, isolated Docker containers to prevent malicious LLM code generation from executing arbitrary operations on the host SRE server. - Principle of Least Privilege: GitHub Personal Access Tokens (PATs) used by the
pr_authoragent are strictly scoped to therepo:writecapability and are designed for rapid rotation. - Air-Gapped Telemetry: No user source code is ever transmitted via SSE payloads; the system relies heavily on metadata, diff hashes, and rule IDs to stream pipeline progress.
- Secure Secret Transport: API tokens and GitHub credentials are never passed via command-line arguments (which are visible in process logs) and are injected using
http.extraheaderconfiguration for git operations. - Strict Input Validation: All external inputs, specifically target repository URLs, are scrubbed and validated via strict regex allowlists to thwart Server-Side Request Forgery (SSRF) and command injection vectors prior to cloning.
- Advanced Distributed Tracing: Expanding the
repo_mapperto trace vulnerabilities across runtime microservice boundaries via OpenTelemetry integrations, building on top of the already implemented Multi-Repository wildcard execution engine (github.com/org/*) and ASTKnowledgeGraphservice boundary mappings. - IDE Integration: Exposing the FastAPI backend via a Language Server Protocol (LSP) to allow the LangGraph pipeline to operate directly within VSCode/JetBrains environments.
- Advanced Self-Healing via MCTS: Implementing Monte Carlo Tree Search (MCTS) within the
validatorloop to explore multiple repair pathways simultaneously, testing competing branches and selecting the one with the highest terminal confidence score.