Skip to content

Latest commit

Β 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸš€ RAG-Cratoss

Retrieval-Augmented Generation Pipeline for IoT & Network Security Documents

A complete, end-to-end RAG system that ingests IoT/network security PDFs, chunks and embeds them using Sentence Transformers, stores vectors in ChromaDB, retrieves context via hybrid retrieval (BM25 + semantic search with RRF fusion), reranks with a cross-encoder, applies a 3-tier confidence check, and generates grounded answers using Meta's Llama 3.2 running locally via Ollama β€” all without any API keys or cloud dependencies.


🎯 What Is This Project?

RAG-Cratoss is a domain-specific question-answering system built on the Retrieval-Augmented Generation (RAG) architecture. Instead of relying on a general-purpose LLM that may hallucinate, it:

  1. Ingests your own PDF documents (IoT specs, protocol RFCs, hardware datasheets)
  2. Splits them into semantically meaningful chunks
  3. Embeds each chunk into a 384-dimensional vector space
  4. Stores them in a persistent vector database (ChromaDB)
  5. Retrieves the most relevant chunks using hybrid search (BM25 + semantic + RRF fusion)
  6. Reranks candidates with a cross-encoder for precision
  7. Evaluates confidence through a 3-tier system (none / low / full)
  8. Generates an accurate, grounded answer using the retrieved context, with an IoT-specific fallback to general knowledge if the database is missing technical data.

πŸ’‘ Key Insight: The system prioritizes your documents. If a question is about IoT/hardware but the answer isn't in your files, it uses the LLM's expert knowledge (clearly tagged). For non-technical/off-topic questions not in the context, it strictly refuses to answer.


🧠 How It Works

                         πŸ“„ Your PDFs (IoT, Protocols, Hardware)
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 1: PDF Loading       β”‚
                         β”‚   (loader.py β€” PyPDF)        β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 2: Text Chunking     β”‚
                         β”‚   (chunker.py β€” 1000 chars)  β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 3: Embedding         β”‚
                         β”‚   (embedder.py β€” MiniLM)     β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 4: Vector Storage    β”‚
                         β”‚   (ChromaDB β€” 577 chunks)    β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
         User Question ──────────────────
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 5: Hybrid Retrieval  β”‚
                         β”‚   BM25 + Semantic + RRF      β”‚
                         β”‚   (hybrid_retriever.py)      β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 6: Cross-Encoder     β”‚
                         β”‚   Reranking (reranker.py)    β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 7: Confidence Check  β”‚
                         β”‚   3-tier (none/low/full)     β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β–Ό           β–Ό                 β–Ό
                  NONE        LOW              FULL
              (score<0.25) (0.25–0.45)       (β‰₯0.45)
                  β”‚           β”‚                 β”‚
                  β–Ό           β–Ό                 β–Ό
              ❌ Reject   ⚠️ LLM +          βœ… LLM
              fallback    warning prefix     full answer
                                        β”‚
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Phase 8: LLM Generation    β”‚
                         β”‚   (Llama 3.2 via Ollama)     β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                        β”‚
                                   πŸ’¬ Answer
                              (with source citations)

πŸ“‚ Project Structure

RAG-Cratoss/
β”œβ”€β”€ main.py                    # πŸ†• Interactive Q&A loop (Primary entry point)
β”œβ”€β”€ data/
β”‚   └── pdfs/
β”‚       β”œβ”€β”€ architecture/          # IoT architecture docs (NIST SP 800-183)
β”‚       β”œβ”€β”€ hardware/              # Hardware datasheets (Arduino UNO R3)
β”‚       β”œβ”€β”€ protocols/             # Protocol specs (MQTT, CoAP RFC 7252)
β”‚       └── security/              # Security docs (IoT device pentesting)
β”œβ”€β”€ ingestion/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ loader.py                  # Phase 1 β€” PDF loading with metadata
β”‚   β”œβ”€β”€ chunker.py                 # Phase 2 β€” Recursive text chunking
β”‚   └── embedder.py                # Phase 3 β€” Embedding & ChromaDB storage
β”œβ”€β”€ rag/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ retriever.py               # Semantic similarity retrieval (original)
β”‚   β”œβ”€β”€ hybrid_retriever.py        # πŸ†• Hybrid retrieval (BM25 + Semantic + RRF)
β”‚   β”œβ”€β”€ reranker.py                # πŸ†• Cross-encoder reranking
β”‚   └── pipeline.py                # Full RAG pipeline (hybrid β†’ rerank β†’ confidence β†’ generate)
β”œβ”€β”€ vectorstore/
β”‚   └── chroma_db/                 # Persisted ChromaDB vector store (~600 chunks)
β”œβ”€β”€ api/
β”‚   └── main.py                    # FastAPI endpoint (planned)
β”œβ”€β”€ .env                           # Environment variables (optional)
β”œβ”€β”€ requirements.txt               # Python dependencies
└── README.md

βš™οΈ Tech Stack

Component Technology Details
Language Python 3.11+
Framework LangChain Orchestration & chain building
PDF Parsing PyPDF Page-by-page PDF extraction
Text Splitting RecursiveCharacterTextSplitter 1000-char chunks with 200-char overlap
Embedding Model sentence-transformers/all-MiniLM-L6-v2 384-dimensional, normalized embeddings
Vector Store ChromaDB (persistent) Local storage, cosine similarity
Keyword Search rank_bm25 (BM25Okapi) πŸ†• Sparse keyword retrieval
Reranker cross-encoder/ms-marco-MiniLM-L-6-v2 πŸ†• Cross-encoder reranking (local)
LLM Meta Llama 3.2 (3B) Local expert with IoT fallback logic
LLM Runtime Ollama Local inference, GPU accelerated
API FastAPI + Uvicorn REST endpoint (planned)

πŸ”§ Setup & Installation

Prerequisites

  • Python 3.11+
  • Ollama (for running Llama 3.2 locally)
  • NVIDIA GPU (optional but recommended β€” e.g., RTX 4060)

Step 1: Clone the Repository

git clone https://github.com/your-username/RAG-Cratoss.git
cd RAG-Cratoss

Step 2: Create Virtual Environment & Install Dependencies

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Step 3: Install Ollama & Download Llama 3.2

# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Download Llama 3.2 (3B model, ~2GB)
ollama pull llama3.2

Step 4 (Optional): GPU Setup for WSL2

If you have an NVIDIA GPU and are using WSL2:

# Verify GPU is visible in WSL
nvidia-smi

# Restart Ollama to detect GPU
sudo systemctl restart ollama

# Verify GPU is being used
ollama ps

πŸ’‘ Performance: With an RTX 4060, responses take 5-10 seconds. On CPU, expect 1-3 minutes per query.


πŸš€ Usage

Run the Full Pipeline

# Make sure Ollama is running
ollama serve &

# Start the interactive Q&A session
python main.py

This will:

  1. Load the ChromaDB vector store (~600 embedded document chunks)
  2. Initialize Llama 3.2 locally via Ollama
  3. Start a continuous loop where you can type your questions interactively
  4. Distinguish between document-based answers and general IoT knowledge fallback

Run Individual Phases

# Phase 1: Load PDFs
python -m ingestion.loader

# Phase 2: Chunk documents
python -m ingestion.chunker

# Phase 3: Embed & store in ChromaDB
python -m ingestion.embedder

# Phase 4: Test retrieval only
python -m rag.retriever

# Phase 5: Full RAG pipeline (retrieve + generate)
python rag/pipeline.py

πŸ“‹ Pipeline Phases in Detail

Phase 1 β€” PDF Loading (ingestion/loader.py)

Recursively loads all PDF files from data/pdfs/ and returns LangChain Document objects with enriched metadata.

  • Walks through architecture/, hardware/, and protocols/ subdirectories
  • Loads each PDF page-by-page using PyPDFLoader
  • Tags each page with a category derived from the subfolder name
  • Adds file_name to metadata for traceability

Phase 2 β€” Text Chunking (ingestion/chunker.py)

Splits loaded pages into smaller, overlapping chunks optimized for embedding and retrieval.

Parameter Value
chunk_size 1000 characters
chunk_overlap 200 characters
separators \n\n β†’ \n β†’ . β†’ β†’ ""
  • Uses RecursiveCharacterTextSplitter to keep paragraphs and sentences intact
  • Preserves all original metadata (source, page, category, file_name)
  • Adds a chunk_index to each chunk for traceability

Phase 3 β€” Embedding & Storage (ingestion/embedder.py)

Generates embeddings using Sentence Transformers and persists them in ChromaDB.

  • Embedding model: sentence-transformers/all-MiniLM-L6-v2 (384-dimensional)
  • Processes chunks in batches of 50 to avoid memory issues
  • Persists the vector store to vectorstore/chroma_db/
  • Collection name: rag_cratoss_docs
  • Result: ~600 document chunks embedded and stored

Phase 4 β€” Semantic Retrieval (rag/retriever.py)

The original retriever. Loads the persisted ChromaDB vector store and retrieves the most semantically relevant chunks using cosine similarity.

Retrieval Modes:

Mode Method Description
Basic retrieve(query) Top-K similarity search
Scored retrieve_with_scores(query) Returns relevance scores alongside results
Filtered retrieve_with_filter(query, filter) Metadata-filtered search (e.g., by category)
LangChain get_langchain_retriever() Returns a retriever for LangChain pipelines

πŸ’‘ This module is still available for standalone use, but pipeline.py now uses HybridRetriever instead.


Phase 5 β€” πŸ†• Hybrid Retrieval (rag/hybrid_retriever.py)

Combines BM25 keyword search with ChromaDB semantic search and merges them using Reciprocal Rank Fusion (RRF).

How it works:

  1. At init, loads ALL document chunks from ChromaDB and builds a BM25 index (no re-ingestion needed)
  2. At query time, runs both BM25 and semantic search independently
  3. Merges both ranked lists using RRF: score(d) = Ξ£ 1/(k + rank_i(d)) with k=60
  4. Returns top-K fused results sorted by descending RRF score
Method Signature Returns
retrieve_hybrid (query, top_k=5) list[(fused_score, doc_text, metadata)]

Why hybrid? BM25 excels at exact keyword matches (e.g., "MQTT", "CoAP") while semantic search captures meaning. RRF combines the best of both.


Phase 6 β€” πŸ†• Cross-Encoder Reranking (rag/reranker.py)

Refines the hybrid results using a cross-encoder that scores each (query, document) pair jointly through a single transformer pass.

Parameter Value
Model cross-encoder/ms-marco-MiniLM-L-6-v2
Device Auto-detect (CUDA if available, else CPU)
Default top_n 3
Method Signature Returns
rerank (query, chunks, top_n=3) list[(ce_score, doc_text, metadata)]

Why rerank? Bi-encoder retrieval scores query and document independently. A cross-encoder processes them together, capturing deeper token-level interactions β€” dramatically improving precision on the small candidate set.


Phase 7 β€” πŸ†• 3-Tier Confidence Check (pipeline.py)

After reranking, the pipeline evaluates the top chunk's score to determine how to respond:

Tier Condition Behavior
NONE max_score < -10.0 πŸ€– IoT Fallback: Use LLM internal knowledge ONLY for IoT topics
LOW -10.0 ≀ max_score < 1.0 ⚠️ Call LLM with weakly matched context
FULL max_score β‰₯ 1.0 βœ… Call LLM with strong context and citations

The helper function get_confidence_tier(max_score) returns "none", "low", or "full".

Every query logs which tier was triggered:

🎯 Confidence tier: LOW (max score: 0.3800)

Phase 8 β€” Generation Pipeline (rag/pipeline.py)

The core orchestration layer that connects hybrid retrieval β†’ reranking β†’ confidence check β†’ LLM generation.

Pipeline Flow:

User Question
     β”‚
     β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Hybrid Retrieve  β”‚  ← BM25 + Semantic + RRF fusion
β”‚ (top_k=5)        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Cross-Encoder    β”‚  ← Rerank 5 β†’ 3 chunks
β”‚ Rerank (top_n=3) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β”‚
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Confidence Tier  β”‚  ← Check max reranked score
β””β”€β”€β”¬β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”˜
   β”‚     β”‚     β”‚
   β–Ό     β–Ό     β–Ό
 NONE   LOW   FULL
   β”‚     β”‚     β”‚
   β–Ό     β–Ό     β–Ό
 πŸ€–    ⚠️+πŸ€–  πŸ€–
IoT    Weak  Full
Fallback Context Answer

Key Configuration:

Parameter Value Description
LLM_MODEL llama3.2 Meta's Llama 3.2 3B via Ollama
MAX_NEW_TOKENS 512 Maximum tokens in generated answer
TEMPERATURE 0.3 Low temperature for factual answers
RELEVANCE_THRESHOLD -10.0 Tier boundary: NONE vs LOW
LOW_CONFIDENCE_THRESHOLD 1.0 πŸ†• Tier boundary: LOW vs FULL
TOP_K 5 Chunks retrieved by hybrid retriever
RERAN_TOP_N 3 πŸ†• Chunks kept after reranking

Prompt Strategy: The LLM is instructed to act as an IoT/network security expert and answer only from the provided context. If the context is insufficient, it explicitly says so β€” preventing hallucination.


πŸ“Š Retrieval Quality & Score Guide

Confidence Tiers (Post-Reranking)

Tier Score Range Meaning Pipeline Action
FULL β‰₯ 1.0 🟒 High confidence β€” strong match βœ… Context Answer
LOW -10.0 – 1.0 🟑 Moderate β€” weakly matched ⚠️ Weakly Matched Answer
NONE < -10.0 πŸ€– No match found πŸ€– IoT Fallback vs Refusal

Example Results

In-domain query: "What is MQTT protocol?"

Result Source File Score Tier
#1 mqtt_protocol_spec.pdf (Page 0) 🟒 4.2500 FULL
#3 mqtt_protocol_spec.pdf (Page 6) 🟒 2.6800 FULL

βœ… All results correctly from the MQTT spec. Tier: FULL β€” generates confident answer.

IoT Fallback Query: "How to blink an LED on ESP32?"

Result Source File Score Tier
#1 arduino_uno_datasheet.pdf 🟑 -8.4500 LOW

πŸ€– Tier: LOW/NONE β€” Since the topic is IoT but the context is weak, the LLM uses its General Knowledge Fallback to provide the code.

Out-of-domain query: "What is the capital of India?"

Result Source File Score Tier
#1 coap_protocol_rfc7252.pdf (Page 3) πŸ”΄ -11.0200 NONE

❌ Tier: NONE β€” Outside IoT expertise and no context found. Pipeline refuses to answer to prevent hallucinations.


πŸ“š Knowledge Base Documents

The system currently has ~600 embedded chunks from these documents:

Category Document Description
Architecture nist_iot_architecture.pdf NIST SP 800-183 β€” Networks of 'Things'
Hardware arduino_uno_datasheet.pdf Arduino UNO R3 full datasheet
Protocols mqtt_protocol_spec.pdf MQTT V3.1 Protocol Specification
Protocols coap_protocol_rfc7252.pdf CoAP β€” RFC 7252 (Constrained Application Protocol)

πŸ’‘ Add your own PDFs: Simply drop PDF files into data/pdfs/<category>/ and re-run the ingestion pipeline (Phases 1-3).


πŸ› οΈ Configuration

All pipeline parameters can be tuned in the respective files:

rag/pipeline.py

LLM_MODEL = "llama3.2"              # Change to "llama3.2:1b" for faster CPU inference
MAX_NEW_TOKENS = 512                # Max response length
TEMPERATURE = 0.3                   # Higher = more creative, Lower = more factual
RELEVANCE_THRESHOLD = -10.0          # Tier boundary: NONE vs LOW
LOW_CONFIDENCE_THRESHOLD = 1.0      # πŸ†• Tier boundary: LOW vs FULL
TOP_K = 5                           # Number of chunks to retrieve
RERAN_TOP_N = 3                     # πŸ†• Chunks kept after reranking

ingestion/chunker.py

CHUNK_SIZE = 1000               # Characters per chunk
CHUNK_OVERLAP = 200             # Overlap between consecutive chunks

πŸ—ΊοΈ Roadmap

  • Phase 1 β€” PDF Loading with metadata enrichment
  • Phase 2 β€” Recursive text chunking
  • Phase 3 β€” Sentence Transformer embeddings + ChromaDB storage
  • Phase 4 β€” Semantic retrieval with relevance scoring
  • Phase 5 β€” πŸ†• Hybrid retrieval (BM25 + Semantic + RRF fusion)
  • Phase 6 β€” πŸ†• Cross-encoder reranking
  • Phase 7 β€” πŸ†• 3-tier confidence-aware response system
  • Phase 8 β€” LLM generation pipeline (Llama 3.2 via Ollama)
  • Phase 9 β€” FastAPI REST endpoint (api/main.py)
  • Phase 10 β€” Web UI / Chat interface
  • Phase 11 β€” Evaluation metrics (BLEU, ROUGE, faithfulness)
  • Phase 12 β€” Multi-document conversation memory

πŸ“Š System Requirements

Component Minimum Recommended
RAM 8 GB 16 GB
Storage 5 GB 10 GB
GPU None (CPU works) NVIDIA RTX 4060+ (8GB VRAM)
OS Ubuntu 20.04 / WSL2 Ubuntu 22.04 / WSL2
Python 3.10 3.11+

πŸ“ License

This project is for educational and research purposes.


πŸ‘€ Author

Dharshan Kumar J

Built as a hands-on project to learn and implement Retrieval-Augmented Generation from scratch.

About

RAG-Cratoss is a domain-specific question-answering system built on the Retrieval-Augmented Generation (RAG) architecture

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages