Skip to content

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

applied-ai-lab

A working notebook for senior-backend-engineer-meets-applied-AI: production-shaped patterns for LLM apps, RAG, MCP, and agentic systems. Code-first, diagrams where they help, every component explained from first principles.

Status: complete — six concept explainers, a set of hands-on builds, and a fully dockerized capstone.


Why this repo exists

Most "AI engineering" content is one of two extremes — Hello-World notebooks that don't survive contact with production, or framework demos (pip install langchain) that hide the mechanics. This repo sits in the middle:

  • Built from primitives first — raw HTTP / official SDK / hand-written loops — then shows what a framework would add
  • Documented in depth — every component has a "how does this actually work?" doc next to it
  • Architecture-aware — chunking strategies, retrieval eval, agent loops, observability, failure modes
  • Reused, not invented — the same pieces (Kafka, microservices, OTel) you'd use for non-AI distributed systems, applied here

Repo map

Path What lives here
docs/01-concepts/ Six architecture-depth explainers w/ Mermaid diagrams (LLM internals, embeddings, RAG, agent loop, MCP, observability)
src/01-llm-basics/ Plain LLM call → streaming → tool_use atom. Understand the request/response shape.
src/02-rag/ RAG pipeline from scratch + eval harness w/ Recall@k, MRR. Real numbers.
src/03-agent/ Tool-using agent loop + hand-rolled MCP server (raw JSON-RPC over stdio) + agent that uses MCP as tool transport
src/04-capstone/ Building Ops Assistant — FastAPI + agent + MCP + RAG + Kafka ingest + OpenTelemetry → Jaeger, all in docker-compose

The capstone — Building Ops Assistant

A production-shaped agent that mirrors real enterprise-AI architecture:

flowchart LR
    User[Browser / curl] -->|POST /chat| API[FastAPI<br/>SSE stream]
    API --> Agent[Agent loop<br/>cost cap + cache]
    Agent <-->|messages + tool_use| Claude[Anthropic API]
    Agent --> MCPClient[MCPClient<br/>JSON-RPC stdio]
    MCPClient <--> Server[MCP server<br/>subprocess]
    Server --> Telemetry[query_telemetry]
    Server --> SearchManual[search_manual]
    Server --> WorkOrder[create_work_order<br/>PENDING_APPROVAL]

    Producer[Synthetic sensor<br/>producer] -->|events| Kafka[(Kafka<br/>bms.sensor.events)]
    Kafka --> Ingest[Ingest worker]
    Ingest --> Store[(SQLite<br/>telemetry store)]
    Telemetry --> Store
    SearchManual --> Chroma[(Chroma<br/>vector store)]
    WorkOrder -.->|approval card| UI[Browser UI]

    Agent -.->|spans| OTel[OpenTelemetry]
    MCPClient -.-> OTel
    OTel --> Jaeger[(Jaeger)]

    classDef ai fill:#e8f4ff,stroke:#0366d6
    classDef mcp fill:#fef3c7,stroke:#f59e0b
    classDef obs fill:#e0f2fe,stroke:#0284c7
    classDef data fill:#dcfce7,stroke:#16a34a
    classDef stream fill:#fde2e2,stroke:#ef4444
    class Claude ai
    class MCPClient,Server mcp
    class OTel,Jaeger obs
    class Chroma,Telemetry,SearchManual,Store data
    class Producer,Kafka,Ingest stream
Loading

Eight production-shaping decisions (see src/04-capstone/README.md):

  1. Event-driven ingest (Kafka → ingest worker → time-series store)
  2. RAG over domain documents w/ grounded prompting + citation validation
  3. Hand-rolled MCP server so the protocol is something implemented, not configured
  4. Agent loop w/ MAX_ITERS, tool-errors-as-data, parallel tool use
  5. Approval gate on side-effecting tools (create_work_order → PENDING_APPROVAL)
  6. Prompt caching on stable system + tool schemas (~10% cost on cache reads)
  7. Per-request cost cap with structured abort
  8. OpenTelemetry GenAI semantic conventions on every LLM + tool call, viewable in Jaeger

Quick start

# clone
git clone https://github.com/XKushal/applied-ai-lab.git
cd applied-ai-lab

# secrets
cp .env.example .env  # add your ANTHROPIC_API_KEY

# install (uv handles venv + deps)
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync

# ingest the RAG corpus once
uv run python src/02-rag/01_ingest.py

# the fast path: local API + UI, spans to console
cd src/04-capstone
uv run uvicorn app.main:app --reload --host 0.0.0.0 --port 8000
# → http://localhost:8000 in your browser

# OR the full path: docker-compose with Kafka + Jaeger
cd src/04-capstone/docker
docker compose up --build
# → http://localhost:8000 for UI, http://localhost:16686 for Jaeger

What's here

  • Concepts — six explainers w/ Mermaid diagrams: LLM internals, embeddings + retrieval, RAG architecture, the agent loop, MCP, observability
  • Hands-on builds — plain LLM calls → RAG from scratch w/ eval → tool-using agent → hand-rolled MCP server
  • Capstone — Building Ops Assistant: FastAPI + agent + MCP + RAG + Kafka ingest + OpenTelemetry, fully dockerized

Built by Kushal Singh — senior backend / distributed systems / applied AI.

About

Senior-engineer-meets-applied-AI: production-shaped patterns for LLM apps, RAG, MCP, and agentic systems.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages