An end-to-end Retrieval-Augmented Generation (RAG) pipeline built entirely with open-source tools. This application allows users to upload PDF documents and ask context-specific questions. All text extraction, embedding, and generation occurs 100% locally on-device, ensuring zero API costs and complete data privacy.
This project strictly follows the standard RAG architecture:
- Data Ingestion:
PyPDFLoaderextracts text from the uploaded PDF. - Chunking:
RecursiveCharacterTextSplitterdivides the text into manageable 1000-character chunks with a 200-character overlap to preserve semantic context. - Embedding: Chunks are vectorized using the local
nomic-embed-textmodel via Ollama. - Vector Storage: Vectors are stored locally in a
Chromadatabase for rapid similarity search. - Retrieval & Generation: User queries are embedded, matched against the vector store using cosine similarity, and passed to a local
Llama 3model alongside the retrieved context to generate deterministic, grounded answers.
- Language: Python
- Orchestration: LangChain
- LLM Engine: Ollama (Llama 3 for generation, Nomic for embeddings)
- Vector Database: ChromaDB
- User Interface: Gradio
You must have Ollama installed on your machine to run the local models.
Once Ollama is installed, pull the required models via your terminal:
ollama run llama3
ollama pull nomic-embed-text