A minimal Retrieval Augmented Generation (RAG) system demonstrating production-grade AI engineering practices.
The goal is not to build a complete regulatory assistant but to demonstrate:
- document ingestion
- embedding pipelines
- vector search
- RAG evaluation
- structured outputs
- observability
- cost awareness
- API-first design
Financial institutions must navigate large regulatory documents.
Users need answers that are:
- accurate
- grounded in source documents
- explainable
- traceable
The system allows users to ask questions about regulatory documents and receive answers with citations.
Minimal Proof of Understanding.
Out of scope:
- authentication
- frontend
- multi-tenancy
- production deployment
| Source | URL | Purpose |
|---|---|---|
| Basel III Framework | https://www.bis.org/basel_framework/ | Regulatory corpus |
| Basel Committee Publications | https://www.bis.org/bcbs/publ/ | Additional guidance |
| ESMA Guidelines | https://www.esma.europa.eu/document-library/guidelines | European regulation |
| SEC Rules and Regulations | https://www.sec.gov/rules-regulations | US regulation |
| MiFID II Overview | https://finance.ec.europa.eu/capital-markets-union-and-financial-markets/financial-markets/securities-markets/mifid-ii-and-mifir_en | Market regulation |
| ID | Requirement |
|---|---|
| FR-1 | Ingest PDF documents |
| FR-2 | Extract and clean text |
| FR-3 | Chunk documents |
| FR-4 | Generate embeddings |
| FR-5 | Store embeddings in vector database |
| FR-6 | Accept questions through REST API |
| FR-7 | Retrieve Top-K chunks |
| FR-8 | Generate grounded answer using OpenAI |
| FR-9 | Return citations |
| FR-10 | Return structured JSON response |
| FR-11 | Expose retrieval diagnostics |
| ID | Requirement | Target |
|---|---|---|
| NFR-1 | Response latency | <10 sec |
| NFR-2 | Retrieval quality | Recall@5 > 80% |
| NFR-3 | Faithfulness | >0.80 |
| NFR-4 | Containerization | Docker |
| NFR-5 | Code quality | Type hints + linting |
| NFR-6 | Observability | LangSmith traces |
| NFR-7 | Cost visibility | Token usage recorded |
| ID | |||
|---|---|---|---|
| US-1 | As a compliance analyst | I want to ask questions in natural language | so that I can locate regulations quickly |
| US-2 | As an auditor | I want citations | so that I can verify answers |
| US-3 | As an engineer | I want retrieval diagnostics | so that I can evaluate system quality |
| US-4 | As a manager | I want cost statistics | so that I can estimate operational expenses |
| ID | Test | Expected Result |
|---|---|---|
| TC-1 | Known answer question | Correct chunk appears in Top-5 |
| TC-2 | Out-of-scope question | "I don't know" response |
| TC-3 | Citation validation | At least one source returned |
| TC-4 | JSON validation | Schema passes |
| TC-5 | Retrieval benchmark | Recall@5 > 80% |
| TC-6 | Evaluation benchmark | Faithfulness > 0.80 |
Retrieval metrics:
- Recall@5: |relevant items in the top 5 positions| / |relevant items for that query|
- MRR (Mean Reciprocal Rank): Expected_value(1/rank_i) where rank_i is the position of the first relevant item in query i.
RAG evaluation metrics:
- Faithfulness: |statements supported by context| / |all the statments in the answer|
- Context Precision: measures whether the most relevant information is ranked at the very top of the retrieved context list produced by RAG
- Answer Relevancy: measures how directly a generated answer addresses the user's initial question
Evaluation frameworks:
| ID | Deliverable |
|---|---|
| DOD-1 | PDF ingestion implemented |
| DOD-2 | Embedding pipeline implemented |
| DOD-3 | Qdrant operational |
| DOD-4 | OpenAI integration completed |
| DOD-5 | Evaluation dataset created |
| DOD-6 | Ragas benchmark executed |
| DOD-7 | DeepEval benchmark executed |
| DOD-8 | Docker image created |
| DOD-9 | CI pipeline operational |
| DOD-10 | README completed |
| Decision | Chosen | Alternative | Reason |
|---|---|---|---|
| LLM | OpenAI GPT-4.1 | Local Llama | Focus on application engineering |
| Vector DB | Qdrant | Pinecone | Open source and reproducible |
| Framework | FastAPI | Flask | Strong typing and OpenAPI support |
| Data Engine | Polars | Pandas | Better scalability |
- OpenAI API
- Prompt engineering
- Structured outputs
- Tool calling
- Embeddings
- Chunking
- Vector search
- Citation generation
- Polars
- DuckDB
- Parquet
- Apache Arrow
- FastAPI
- Pydantic
- SQLAlchemy
- Ragas
- DeepEval
- Golden dataset
- Docker
- GitHub Actions
- Logging
- Metrics
- Prompt injection protection
- Input validation
- Token usage tracking
- Cost dashboard