A privacy-first, open-source alternative to Google's NotebookLM. Query your documents with AI-powered semantic search, source-grounded answers, and automatic citations.
- โ๏ธ Notebook Settings: Customize memory, retrieval count, score thresholds, Self-RAG quality gates, and prompts independently per notebook natively via UI.
- ๐ Document Hub: Manage PDF and Word documents, track real-time hardware status (RAM/VRAM), and configure intelligent settings with hardware-aware warnings.
- ๐ง Self-RAG Pipeline: Multi-hop retrieval with automatic quality scoring (groundedness, relevance, utility) and a repair agent that retries failed searches with a new strategy โ delivering significantly more accurate and grounded answers than a single-pass RAG.
- ๐ค Co-RAG Pipeline: Collaborative GeneratorโReviewer architecture that iteratively refines answers through peer critique โ the Generator drafts a holistic answer, the Reviewer diagnoses gaps and hallucinations, and the loop repeats until the answer is verified or the turn limit is reached.
- โก Dual-Pipeline Comparison: Both Self-RAG and Co-RAG run independently on every query, each with its own chat memory and reformulated context. Results are displayed side-by-side in two tabs for direct comparison.
- ๐ Hybrid Search: Intelligently retrieve relevant document sections using combined semantic search (FAISS embeddings) and BM25 full-text search with configurable weighting.
- ๐ค Grounded AI Responses: Get answers strictly based on your documents with automatic source citations
- ๐ Multi-language Support: Ask questions in Vietnamese, English, or other languages and receive answers in your preferred language
- ๐พ Persistent Storage: All documents, chat history, and notes are saved to a local SQLite database
- ๐ Study Notes: Save important Q&A pairs as notes to build a personalized study guide
- โก Privacy First: Runs entirely locally with Ollama-your documents never leave your machine
- ๐จ NotebookLM-Inspired UI: Clean, intuitive Streamlit interface mirroring the NotebookLM experience
| Component | Technology | Purpose |
|---|---|---|
| UI Framework | Streamlit | Interactive web interface |
| Orchestration | LangChain | RAG pipeline & prompt management |
| Vector Database | FAISS | Fast similarity search on CPU |
| Embeddings | Sentence Transformers | Multi-language text vectorization (paraphrase-multilingual-mpnet-base-v2) |
| LLM | Ollama + Qwen2.5:7b | Local inference with Vietnamese support |
| Database | SQLite | Metadata and chat history persistence |
| Document Processing | PyMuPDF & python-docx | Extract text from PDFs and Word documents |
- Python: 3.8 or higher
- Ollama: Running locally with
qwen2.5:7bmodel pulled - System Memory: 8GB RAM minimum (16GB+ recommended for optimal performance)
- GPU (optional): CUDA-capable GPU for faster embeddings (CPU fallback available)
# Clone the repository
git clone https://github.com/dungtq2k5/smartdoc-ai.git
cd smartdoc-ai
# Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt# Download Ollama from https://ollama.ai/
# Then pull the Qwen2.5 model (required)
ollama pull qwen2.5:7b
# Start Ollama server (runs in background)
ollama serve# In a new terminal, ensure venv is activated
source venv/bin/activate
# Start Streamlit
streamlit run app.pyThe app will open at http://localhost:8501.
- Click "+ Create New Notebook" on the dashboard
- Enter a notebook name (e.g., "Machine Learning Research")
- Optionally add a description
- Select a notebook and open the Notebook Settings panel in the sidebar
- Adjust retrieved chunk count, conversation memory, temperature, or add special System Prompts for this specific notebook
- Click "Apply Settings" to securely persist your configuration
- Select a notebook from the dashboard
- Go to the Source Hub (left sidebar)
- Click "Upload Documents" and select one or more PDF or DOCX files
- Review and confirm uploadsโdocuments are automatically processed
- Select which documents to query using checkboxes
- Type your question in the chat input
- SmartDoc retrieves relevant sections and generates a grounded answer
- Click on retrieved sources beneath each answer to verify citations
- After getting a good answer, click "Save as Note" below the response
- Edit the note title if needed
- Saved notes are compiled in the Notes Panel (right sidebar)
- Remove a document: Click the โฎ menu next to a source โ "Delete"
- Delete a notebook: Click the โฎ menu next to notebook name โ "Delete"
- Clear chat history: Click the โฎ menu next to "Chat" title โ "Delete chat history"
All tunable parameters are centralized in core/configs.py:
# RAG Tuning (adjust for inference quality vs speed)
RAG_RETRIEVAL_K: int = 8 # Top K chunks to retrieve
RAG_RETRIEVAL_SCORE_THRESHOLD: float = 1.0 # Euclidean distance (0.0 to 2.0); Lower = stricter filtering
RAG_MAX_CTX_LEN: int = RAG_RETRIEVAL_K * 1000 # Characters sent to LLM
WEIGHT_SEMANTIC: float = 0.5 # Hybrid search: semantic vs keyword balance
WEIGHT_BM25: float = 0.5 # BM25 weight (keyword match)
# LLM Setup
OLLAMA_BASE_URL = "http://localhost:11434" # Ollama server
LLM_MODEL_NAME = "qwen2.5:7b"
LLM_AVG_TEMP = 0.7 # Base temperature; Self-RAG spreads candidates across [0.1, LLM_AVG_TEMP * 1.5]
# Self-RAG Quality Gates
SELF_RAG_MAX_DEPTH: int = 2 # Max recursive repair hops before fallback
SELF_RAG_CANDIDATES: int = 3 # Diverse answer drafts generated per hop
SELF_RAG_MAX_RETRIES_PER_HOP: int = 2 # Sub-query rewrite retries before skipping
SELF_RAG_THRESHOLD_ISSUP: float = 0.70 # Groundedness gate (answer supported by docs?)
SELF_RAG_THRESHOLD_ISREL: float = 0.70 # Relevance gate (chunks match the query?)
SELF_RAG_THRESHOLD_ISUSE: float = 0.70 # Utility gate (answer satisfies user intent?)
# Co-RAG Collaboration
CO_RAG_MAX_RETRIES: int = 3 # Max GeneratorโReviewer turns (0 = no review)
# ...smartdoc-ai/
โโโ app.py # Main Streamlit entry point
โโโ docs/ # Documentation and diagrams
โโโ core/
โ โโโ configs.py # Centralized configuration parameters
โ โโโ rag.py # Dual-pipeline orchestrator (run_dual_rag)
โ โโโ self_rag.py # Self-RAG orchestration pipeline (6-step multi-hop)
โ โโโ co_rag.py # Co-RAG pipeline (GeneratorโReviewer collaboration)
โ โโโ utils.py # RAG pipeline & utility functions
โโโ db/
โ โโโ setup.py # Database schema initialization
โ โโโ crud.py # SQLite database operations
โโโ middlewares/
โ โโโ db_middleware.py # Input validation & sanitization
โโโ data/
โ โโโ smartdoc.db # SQLite database (auto-created)
โ โโโ vectorstores/ # FAISS indices (organized by notebook/source)
โโโ requirements.txt # Python dependencies
โโโ README.md # This file
โโโ LICENSE # MIT License
โโโ CHANGELOG.md # Version historySolution:
- Verify Ollama is running:
ollama servein a separate terminal - Check Ollama is at
http://localhost:11434:curl http://localhost:11434/api/tags - Restart Streamlit:
streamlit run app.py
Solution:
- The app automatically falls back to CPUโthis is normal on limited GPUs
- Reduce the number of selected documents (use checkboxes)
- Lower
RAG_RETRIEVAL_Kin core/configs.py
Solution:
- First embeddings are slower; subsequent queries cache results
- Qwen2.5:7b performs inference at ~10-15 tokens/sec on CPU (single-threaded)
- For faster speeds, use a GPU or upgrade to a larger GPU memory
Solution:
- Check file size (PDFs > 50MB may time out)
- Verify the PDF is text-based (not image scans)
- Try uploading a smaller section of the PDF first
- Type detection โ file magic bytes determine PDF or DOCX; unsupported types are rejected immediately.
- Duplicate detection โ MD5 hash of file bytes; skip if already uploaded to this notebook.
- Markdown conversion โ PDF pages are converted to Markdown via PyMuPDF
page.get_text("dict")with font-size-ratio heading inference (#/##/###) and span-flag bold/italic. DOCX documents are walked in element order (qn("w:p")/qn("w:tbl")) to produce Markdown headings, lists, and pipe tables. - Cleaning โ
_clean_markdown_text()removes stray page-number lines and collapses excess blank lines. - Markdown-aware chunking โ
MarkdownHeaderTextSplitter(#โ####) splits by document structure; oversized sections fall back toRecursiveCharacterTextSplitter. Heading context is retained inside each chunk. - Embedding โ
paraphrase-multilingual-mpnet-base-v2(768-dim), GPU-first with CPU fallback. - Storage โ per-source FAISS
IndexFlatL2index saved to./data/vectorstores/{notebook_id}/{source_id}/; metadata persisted to SQLitesourcestable.
Every user query runs both pipelines independently and returns a combined result displayed in two tabs.
Self-RAG (Vertical โ multi-hop):
- Intent routing: Classify query as greeting (skip retrieval) or factual (proceed). Layer 1 uses regex; Layer 2 uses an LLM call for ambiguous inputs.
- Query reformulation: Rewrite follow-up questions into standalone queries using Self-RAG's own chat history.
- Search planning: LLM decomposes the query into 1โ3 independent sub-queries covering different aspects.
- Hybrid retrieval + retry: Run FAISS semantic + BM25 keyword search per sub-query; cross-encoder re-ranks results. Failed sub-queries are rewritten and retried up to
SELF_RAG_MAX_RETRIES_PER_HOPtimes. - Candidate generation: Generate
SELF_RAG_CANDIDATESdiverse answer drafts at varied temperatures. - Quality scoring: An LLM judge scores each draft on groundedness (ISSUP), relevance (ISREL), and utility (ISUSE).
- Threshold gate: Accept the best-scoring candidate if all three scores meet configured thresholds.
- Repair & retry: If no candidate passes, a repair agent diagnoses the failure, generates a new search strategy, and retries the full pipeline โ up to
SELF_RAG_MAX_DEPTHrecursive hops. - Stream best answer with source citations; persist confidence score and reasoning trace.
Co-RAG (Horizontal โ iterative peer review):
- Query reformulation: Rewrite follow-up questions into standalone queries using Co-RAG's own isolated chat history.
- Intent routing: Same two-layer greeting detection as Self-RAG, using Co-RAG's own history.
- Holistic retrieval: Single broad FAISS + BM25 pass with cross-encoder re-ranking โ no sub-query decomposition.
- Initial generation (Mode A): Generator LLM drafts a comprehensive answer grounded in retrieved context.
- GeneratorโReviewer loop: Reviewer LLM diagnoses gaps, hallucinations, and contradictions; Generator applies targeted fixes (Mode B). Repeats until
[STATUS: VERIFIED]orCO_RAG_MAX_RETRIESturns exhausted. - Stream final verified answer with sources and full critique trace.
โ Privacy: Your documents never leave your machine โ Offline: Works without internet connectivity โ Cost-Free: No API charges or subscriptions โ Customizable: Full control over embeddings, LLM, and parameters
Version: 1.2.0 (Dual-Pipeline Release) Completion: 87.2% (102/117 development tasks completed) Last Updated: April 20, 2026
- CHANGELOG.md โ Version history and release notes
Detailed architecture diagrams for every component of the dual-pipeline RAG system, each with a Mermaid flowchart and a written step-by-step description.
| Diagram | Description |
|---|---|
| Ingestion Pipeline | Full document ingestion flow: file-type detection, duplicate check, PDF/DOCXโMarkdown conversion, Markdown-aware chunking, embedding, FAISS index creation, SQLite persistence |
| Contextual Query Reformulation | How follow-up queries are rewritten into standalone queries using per-pipeline chat history |
| Greeting Detection | 2-layer intent routing: regex fast-path (Layer 1) + LLM classifier fallback (Layer 2) |
| Search Engine | Pure semantic, pure BM25, and hybrid RRF search modes with proportional k-split and fallback logic |
| Cross-Encoder Reranking | Deduplication for hybrid results, Cross-Encoder scoring, and top-k selection |
| Self-RAG Pipeline | Full 5-step multi-hop pipeline: search plan โ retrieval with surgical retry โ candidate generation โ quality scoring โ repair loop |
| Co-RAG Pipeline | Full holistic pipeline: reformulation โ single-shot retrieval โ Generator Mode A โ GeneratorโReviewer review loop |
| Dual-Pipeline System | End-to-end view of run_dual_rag(): how Self-RAG and Co-RAG run sequentially and display results side-by-side |
This project was developed as part of the Open Source Software Development (OSSD) course at ฤแบกi hแปc Sร i Gรฒn (Saigon University), Spring 2026.
Learning Objectives:
- Build a full-stack AI application with modern Python frameworks
- Implement RAG principles for grounded AI responses
- Work with vector databases and embeddings
- Deploy privacy-first ML systems using local LLMs
- Practice collaborative open-source development
This project is licensed under the MIT Licenseโsee LICENSE for details.
You are free to use, modify, and distribute this software for educational, commercial, or personal purposes.
This project is designed as a learning implementation for the OSSD course and is not actively maintained for external contributions.
Feel free to fork! You can take this code and make it your own:
- Extend it: Add features like Excel document support, OCR for scanned PDFs, or cloud deployment
- Customize it: Modify the UI, change the LLM model, or integrate different vector databases
- Learn from it: Use this as a reference for building your own RAG systems
- Share improvements: If you make significant improvements, feel free to share them as a reference
This codebase is open under the MIT License, so you have full freedom to use and modify it for any purpose.
- Google NotebookLM โ Inspiration for the UI/UX design
- LangChain โ RAG orchestration framework
- Ollama โ Local LLM infrastructure
- FAISS โ Efficient vector search
- Streamlit โ Interactive web framework
- Qwen Team โ Excellent multilingual LLM
If you encounter bugs or have questions, please check the Issues tab first. Feel free to open a new issue if you find a bug.
Happy documenting! ๐โจ