Retrieval-Augmented Generation Pipeline for IoT & Network Security Documents
A complete, end-to-end RAG system that ingests IoT/network security PDFs, chunks and embeds them using Sentence Transformers, stores vectors in ChromaDB, retrieves context via hybrid retrieval (BM25 + semantic search with RRF fusion), reranks with a cross-encoder, applies a 3-tier confidence check, and generates grounded answers using Meta's Llama 3.2 running locally via Ollama β all without any API keys or cloud dependencies.
RAG-Cratoss is a domain-specific question-answering system built on the Retrieval-Augmented Generation (RAG) architecture. Instead of relying on a general-purpose LLM that may hallucinate, it:
- Ingests your own PDF documents (IoT specs, protocol RFCs, hardware datasheets)
- Splits them into semantically meaningful chunks
- Embeds each chunk into a 384-dimensional vector space
- Stores them in a persistent vector database (ChromaDB)
- Retrieves the most relevant chunks using hybrid search (BM25 + semantic + RRF fusion)
- Reranks candidates with a cross-encoder for precision
- Evaluates confidence through a 3-tier system (none / low / full)
- Generates an accurate, grounded answer using the retrieved context, with an IoT-specific fallback to general knowledge if the database is missing technical data.
π‘ Key Insight: The system prioritizes your documents. If a question is about IoT/hardware but the answer isn't in your files, it uses the LLM's expert knowledge (clearly tagged). For non-technical/off-topic questions not in the context, it strictly refuses to answer.
π Your PDFs (IoT, Protocols, Hardware)
β
ββββββββββββββββΌβββββββββββββββ
β Phase 1: PDF Loading β
β (loader.py β PyPDF) β
ββββββββββββββββ¬βββββββββββββββ
β
ββββββββββββββββΌβββββββββββββββ
β Phase 2: Text Chunking β
β (chunker.py β 1000 chars) β
ββββββββββββββββ¬βββββββββββββββ
β
ββββββββββββββββΌβββββββββββββββ
β Phase 3: Embedding β
β (embedder.py β MiniLM) β
ββββββββββββββββ¬βββββββββββββββ
β
ββββββββββββββββΌβββββββββββββββ
β Phase 4: Vector Storage β
β (ChromaDB β 577 chunks) β
ββββββββββββββββ¬βββββββββββββββ
β
User Question ββββββββββββββββββ€
β
ββββββββββββββββΌβββββββββββββββ
β Phase 5: Hybrid Retrieval β
β BM25 + Semantic + RRF β
β (hybrid_retriever.py) β
ββββββββββββββββ¬βββββββββββββββ
β
ββββββββββββββββΌβββββββββββββββ
β Phase 6: Cross-Encoder β
β Reranking (reranker.py) β
ββββββββββββββββ¬βββββββββββββββ
β
ββββββββββββββββΌβββββββββββββββ
β Phase 7: Confidence Check β
β 3-tier (none/low/full) β
ββββββββββββββββ¬βββββββββββββββ
β
βββββββββββββ¬βββββββββ΄βββββββββ
βΌ βΌ βΌ
NONE LOW FULL
(score<0.25) (0.25β0.45) (β₯0.45)
β β β
βΌ βΌ βΌ
β Reject β οΈ LLM + β
LLM
fallback warning prefix full answer
β
ββββββββββββββββΌβββββββββββββββ
β Phase 8: LLM Generation β
β (Llama 3.2 via Ollama) β
ββββββββββββββββ¬βββββββββββββββ
β
π¬ Answer
(with source citations)
RAG-Cratoss/
βββ main.py # π Interactive Q&A loop (Primary entry point)
βββ data/
β βββ pdfs/
β βββ architecture/ # IoT architecture docs (NIST SP 800-183)
β βββ hardware/ # Hardware datasheets (Arduino UNO R3)
β βββ protocols/ # Protocol specs (MQTT, CoAP RFC 7252)
β βββ security/ # Security docs (IoT device pentesting)
βββ ingestion/
β βββ __init__.py
β βββ loader.py # Phase 1 β PDF loading with metadata
β βββ chunker.py # Phase 2 β Recursive text chunking
β βββ embedder.py # Phase 3 β Embedding & ChromaDB storage
βββ rag/
β βββ __init__.py
β βββ retriever.py # Semantic similarity retrieval (original)
β βββ hybrid_retriever.py # π Hybrid retrieval (BM25 + Semantic + RRF)
β βββ reranker.py # π Cross-encoder reranking
β βββ pipeline.py # Full RAG pipeline (hybrid β rerank β confidence β generate)
βββ vectorstore/
β βββ chroma_db/ # Persisted ChromaDB vector store (~600 chunks)
βββ api/
β βββ main.py # FastAPI endpoint (planned)
βββ .env # Environment variables (optional)
βββ requirements.txt # Python dependencies
βββ README.md
| Component | Technology | Details |
|---|---|---|
| Language | Python 3.11+ | |
| Framework | LangChain | Orchestration & chain building |
| PDF Parsing | PyPDF | Page-by-page PDF extraction |
| Text Splitting | RecursiveCharacterTextSplitter | 1000-char chunks with 200-char overlap |
| Embedding Model | sentence-transformers/all-MiniLM-L6-v2 |
384-dimensional, normalized embeddings |
| Vector Store | ChromaDB (persistent) | Local storage, cosine similarity |
| Keyword Search | rank_bm25 (BM25Okapi) |
π Sparse keyword retrieval |
| Reranker | cross-encoder/ms-marco-MiniLM-L-6-v2 |
π Cross-encoder reranking (local) |
| LLM | Meta Llama 3.2 (3B) | Local expert with IoT fallback logic |
| LLM Runtime | Ollama | Local inference, GPU accelerated |
| API | FastAPI + Uvicorn | REST endpoint (planned) |
- Python 3.11+
- Ollama (for running Llama 3.2 locally)
- NVIDIA GPU (optional but recommended β e.g., RTX 4060)
git clone https://github.com/your-username/RAG-Cratoss.git
cd RAG-Cratosspython3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Download Llama 3.2 (3B model, ~2GB)
ollama pull llama3.2If you have an NVIDIA GPU and are using WSL2:
# Verify GPU is visible in WSL
nvidia-smi
# Restart Ollama to detect GPU
sudo systemctl restart ollama
# Verify GPU is being used
ollama psπ‘ Performance: With an RTX 4060, responses take 5-10 seconds. On CPU, expect 1-3 minutes per query.
# Make sure Ollama is running
ollama serve &
# Start the interactive Q&A session
python main.pyThis will:
- Load the ChromaDB vector store (~600 embedded document chunks)
- Initialize Llama 3.2 locally via Ollama
- Start a continuous loop where you can type your questions interactively
- Distinguish between document-based answers and general IoT knowledge fallback
# Phase 1: Load PDFs
python -m ingestion.loader
# Phase 2: Chunk documents
python -m ingestion.chunker
# Phase 3: Embed & store in ChromaDB
python -m ingestion.embedder
# Phase 4: Test retrieval only
python -m rag.retriever
# Phase 5: Full RAG pipeline (retrieve + generate)
python rag/pipeline.pyRecursively loads all PDF files from data/pdfs/ and returns LangChain Document objects with enriched metadata.
- Walks through
architecture/,hardware/, andprotocols/subdirectories - Loads each PDF page-by-page using
PyPDFLoader - Tags each page with a
categoryderived from the subfolder name - Adds
file_nameto metadata for traceability
Splits loaded pages into smaller, overlapping chunks optimized for embedding and retrieval.
| Parameter | Value |
|---|---|
chunk_size |
1000 characters |
chunk_overlap |
200 characters |
separators |
\n\n β \n β . β β "" |
- Uses
RecursiveCharacterTextSplitterto keep paragraphs and sentences intact - Preserves all original metadata (source, page, category, file_name)
- Adds a
chunk_indexto each chunk for traceability
Generates embeddings using Sentence Transformers and persists them in ChromaDB.
- Embedding model:
sentence-transformers/all-MiniLM-L6-v2(384-dimensional) - Processes chunks in batches of 50 to avoid memory issues
- Persists the vector store to
vectorstore/chroma_db/ - Collection name:
rag_cratoss_docs - Result: ~600 document chunks embedded and stored
The original retriever. Loads the persisted ChromaDB vector store and retrieves the most semantically relevant chunks using cosine similarity.
Retrieval Modes:
| Mode | Method | Description |
|---|---|---|
| Basic | retrieve(query) |
Top-K similarity search |
| Scored | retrieve_with_scores(query) |
Returns relevance scores alongside results |
| Filtered | retrieve_with_filter(query, filter) |
Metadata-filtered search (e.g., by category) |
| LangChain | get_langchain_retriever() |
Returns a retriever for LangChain pipelines |
π‘ This module is still available for standalone use, but
pipeline.pynow usesHybridRetrieverinstead.
Combines BM25 keyword search with ChromaDB semantic search and merges them using Reciprocal Rank Fusion (RRF).
How it works:
- At init, loads ALL document chunks from ChromaDB and builds a BM25 index (no re-ingestion needed)
- At query time, runs both BM25 and semantic search independently
- Merges both ranked lists using RRF:
score(d) = Ξ£ 1/(k + rank_i(d))withk=60 - Returns top-K fused results sorted by descending RRF score
| Method | Signature | Returns |
|---|---|---|
retrieve_hybrid |
(query, top_k=5) |
list[(fused_score, doc_text, metadata)] |
Why hybrid? BM25 excels at exact keyword matches (e.g., "MQTT", "CoAP") while semantic search captures meaning. RRF combines the best of both.
Refines the hybrid results using a cross-encoder that scores each (query, document) pair jointly through a single transformer pass.
| Parameter | Value |
|---|---|
| Model | cross-encoder/ms-marco-MiniLM-L-6-v2 |
| Device | Auto-detect (CUDA if available, else CPU) |
| Default top_n | 3 |
| Method | Signature | Returns |
|---|---|---|
rerank |
(query, chunks, top_n=3) |
list[(ce_score, doc_text, metadata)] |
Why rerank? Bi-encoder retrieval scores query and document independently. A cross-encoder processes them together, capturing deeper token-level interactions β dramatically improving precision on the small candidate set.
After reranking, the pipeline evaluates the top chunk's score to determine how to respond:
| Tier | Condition | Behavior |
|---|---|---|
| NONE | max_score < -10.0 |
π€ IoT Fallback: Use LLM internal knowledge ONLY for IoT topics |
| LOW | -10.0 β€ max_score < 1.0 |
|
| FULL | max_score β₯ 1.0 |
β Call LLM with strong context and citations |
The helper function get_confidence_tier(max_score) returns "none", "low", or "full".
Every query logs which tier was triggered:
π― Confidence tier: LOW (max score: 0.3800)
The core orchestration layer that connects hybrid retrieval β reranking β confidence check β LLM generation.
Pipeline Flow:
User Question
β
βΌ
ββββββββββββββββββββ
β Hybrid Retrieve β β BM25 + Semantic + RRF fusion
β (top_k=5) β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β Cross-Encoder β β Rerank 5 β 3 chunks
β Rerank (top_n=3) β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β Confidence Tier β β Check max reranked score
ββββ¬ββββββ¬ββββββ¬ββββ
β β β
βΌ βΌ βΌ
NONE LOW FULL
β β β
βΌ βΌ βΌ
π€ β οΈ+π€ π€
IoT Weak Full
Fallback Context Answer
Key Configuration:
| Parameter | Value | Description |
|---|---|---|
LLM_MODEL |
llama3.2 |
Meta's Llama 3.2 3B via Ollama |
MAX_NEW_TOKENS |
512 | Maximum tokens in generated answer |
TEMPERATURE |
0.3 | Low temperature for factual answers |
RELEVANCE_THRESHOLD |
-10.0 | Tier boundary: NONE vs LOW |
LOW_CONFIDENCE_THRESHOLD |
1.0 | π Tier boundary: LOW vs FULL |
TOP_K |
5 | Chunks retrieved by hybrid retriever |
RERAN_TOP_N |
3 | π Chunks kept after reranking |
Prompt Strategy: The LLM is instructed to act as an IoT/network security expert and answer only from the provided context. If the context is insufficient, it explicitly says so β preventing hallucination.
| Tier | Score Range | Meaning | Pipeline Action |
|---|---|---|---|
| FULL | β₯ 1.0 | π’ High confidence β strong match | β Context Answer |
| LOW | -10.0 β 1.0 | π‘ Moderate β weakly matched | |
| NONE | < -10.0 | π€ No match found | π€ IoT Fallback vs Refusal |
In-domain query: "What is MQTT protocol?"
| Result | Source File | Score | Tier |
|---|---|---|---|
| #1 | mqtt_protocol_spec.pdf (Page 0) |
π’ 4.2500 | FULL |
| #3 | mqtt_protocol_spec.pdf (Page 6) |
π’ 2.6800 | FULL |
β All results correctly from the MQTT spec. Tier: FULL β generates confident answer.
IoT Fallback Query: "How to blink an LED on ESP32?"
| Result | Source File | Score | Tier |
|---|---|---|---|
| #1 | arduino_uno_datasheet.pdf |
π‘ -8.4500 | LOW |
π€ Tier: LOW/NONE β Since the topic is IoT but the context is weak, the LLM uses its General Knowledge Fallback to provide the code.
Out-of-domain query: "What is the capital of India?"
| Result | Source File | Score | Tier |
|---|---|---|---|
| #1 | coap_protocol_rfc7252.pdf (Page 3) |
π΄ -11.0200 | NONE |
β Tier: NONE β Outside IoT expertise and no context found. Pipeline refuses to answer to prevent hallucinations.
The system currently has ~600 embedded chunks from these documents:
| Category | Document | Description |
|---|---|---|
| Architecture | nist_iot_architecture.pdf |
NIST SP 800-183 β Networks of 'Things' |
| Hardware | arduino_uno_datasheet.pdf |
Arduino UNO R3 full datasheet |
| Protocols | mqtt_protocol_spec.pdf |
MQTT V3.1 Protocol Specification |
| Protocols | coap_protocol_rfc7252.pdf |
CoAP β RFC 7252 (Constrained Application Protocol) |
π‘ Add your own PDFs: Simply drop PDF files into
data/pdfs/<category>/and re-run the ingestion pipeline (Phases 1-3).
All pipeline parameters can be tuned in the respective files:
LLM_MODEL = "llama3.2" # Change to "llama3.2:1b" for faster CPU inference
MAX_NEW_TOKENS = 512 # Max response length
TEMPERATURE = 0.3 # Higher = more creative, Lower = more factual
RELEVANCE_THRESHOLD = -10.0 # Tier boundary: NONE vs LOW
LOW_CONFIDENCE_THRESHOLD = 1.0 # π Tier boundary: LOW vs FULL
TOP_K = 5 # Number of chunks to retrieve
RERAN_TOP_N = 3 # π Chunks kept after rerankingCHUNK_SIZE = 1000 # Characters per chunk
CHUNK_OVERLAP = 200 # Overlap between consecutive chunks- Phase 1 β PDF Loading with metadata enrichment
- Phase 2 β Recursive text chunking
- Phase 3 β Sentence Transformer embeddings + ChromaDB storage
- Phase 4 β Semantic retrieval with relevance scoring
- Phase 5 β π Hybrid retrieval (BM25 + Semantic + RRF fusion)
- Phase 6 β π Cross-encoder reranking
- Phase 7 β π 3-tier confidence-aware response system
- Phase 8 β LLM generation pipeline (Llama 3.2 via Ollama)
- Phase 9 β FastAPI REST endpoint (
api/main.py) - Phase 10 β Web UI / Chat interface
- Phase 11 β Evaluation metrics (BLEU, ROUGE, faithfulness)
- Phase 12 β Multi-document conversation memory
| Component | Minimum | Recommended |
|---|---|---|
| RAM | 8 GB | 16 GB |
| Storage | 5 GB | 10 GB |
| GPU | None (CPU works) | NVIDIA RTX 4060+ (8GB VRAM) |
| OS | Ubuntu 20.04 / WSL2 | Ubuntu 22.04 / WSL2 |
| Python | 3.10 | 3.11+ |
This project is for educational and research purposes.
Dharshan Kumar J
Built as a hands-on project to learn and implement Retrieval-Augmented Generation from scratch.