Upload a PDF and ask it questions. A retrieval-augmented pipeline chunks, embeds and searches the document, then answers from the matching passages.
Adapted from advanced_llm_apps/chat_with_X_tutorials/chat_with_pdf in awesome-llm-apps by Shubham Saboo (Apache-2.0). Packaged, documented and maintained here by Mohd Shoaib; a Flutter client is next.
- Upload any PDF and chat with it in a Streamlit chat UI
- Answers come from retrieved chunks of the document, not from the model's memory
- Three variants: OpenAI GPT-4 (
chat_pdf.py), Llama 3 and Llama 3.2 fully local through Ollama - About 30 lines of Python per variant, a clean reference for RAG
Python · Streamlit · embedchain · OpenAI or Ollama (Llama 3 / 3.2)
git clone https://github.com/Shoaib20786/chat-with-pdf-rag.git
cd chat-with-pdf-rag
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
streamlit run chat_pdf.py # OpenAI key in the sidebar
streamlit run chat_pdf_llama3.2.py # local, needs `ollama pull llama3.2`Config: copy .env.example to .env or paste the keys in the app sidebar. Never commit .env.
RAG quality depends on chunk size and overlap: too small and answers lose context, too large and retrieval returns noise. embedchain hides that behind sane defaults, which is why the app stays at 30 lines, but those defaults are the first thing to tune for real documents.
Demo · maintained · Last updated September 2026
Apache-2.0. Original work © Shubham Saboo (awesome-llm-apps); modifications © 2026 Mohd Shoaib. See LICENSE and NOTICE.