Skip to content

Repository files navigation

DocMind

DocMind lets teams upload documents, web pages, and audio/video files, then ask questions about them using a local or cloud AI model. All data stays on your own server. No third-party cloud handles your documents unless you explicitly configure a cloud LLM provider.

Portfolio Note: This is a public showcase repository. It contains architecture documentation, design decisions, and skeleton code structure. The full production implementation is proprietary. The engineering depth, system design, and problem-solving approach demonstrated here reflect the real project.


Live Demo

Demo credentials — Username: admin · Password: admin123


Table of Contents


Problem Statement

Teams working with sensitive documents — contracts, research papers, internal reports — cannot safely send those documents to public AI services like ChatGPT. But they still need to search, summarize, and ask questions about large volumes of content. Doing this manually is slow and error-prone, especially when the document library grows over time.

The problem is not just privacy. It is also fragmentation: files live in different places, different people have access to different documents, and there is no single place to query all of it.

Solution Overview

DocMind is a self-hosted platform where an admin uploads and manages documents, controls who can see what, and provides a chat interface where users can ask questions that are answered by an AI using only the content of the selected documents. Nothing leaves the server unless the admin configures a cloud provider. The system supports local Ollama models out of the box, with optional cloud LLM providers (Anthropic, OpenAI, Google) configurable through a settings panel.


Architecture Overview

DocMind uses a three-tier architecture: a React frontend served by Nginx, a FastAPI backend handling all business logic and AI orchestration, and a data layer consisting of PostgreSQL for relational data and ChromaDB for vector storage. Ollama runs as a sidecar service for local embedding and inference.

System Diagram

flowchart TD
    User["User / Admin\n(Browser)"]

    subgraph Frontend["Frontend Layer"]
        React["React 18\nSPA"]
        Nginx["Nginx\nReverse Proxy"]
    end

    subgraph Backend["Backend Layer"]
        FastAPI["FastAPI\nREST + SSE"]
        Auth["Auth Service\nJWT + Refresh Tokens"]
        RAG["RAG Pipeline\nRetrieval + Prompt Build"]
        Ingest["Ingestion Service\nChunk + Embed"]
        Review["Review Mode\nStructured Analysis"]
        Sched["APScheduler\nCron Re-index"]
    end

    subgraph Data["Data Layer"]
        PG["PostgreSQL\nUsers, Files, Chats"]
        Chroma["ChromaDB\nVector Store"]
    end

    subgraph AI["AI Layer"]
        Ollama["Ollama\nLocal LLM + Embeddings"]
        Cloud["Cloud LLM\nAnthropic / OpenAI / Google"]
    end

    subgraph Infra["Infrastructure"]
        Docker["Docker Compose\nContainer Orchestration"]
    end

    User --> Nginx --> React
    React --> FastAPI
    FastAPI --> Auth
    FastAPI --> RAG
    FastAPI --> Ingest
    FastAPI --> Review
    FastAPI --> Sched
    RAG --> Chroma
    RAG --> Ollama
    RAG --> Cloud
    Ingest --> Chroma
    Ingest --> Ollama
    FastAPI --> PG
    Docker -.-> Frontend & Backend & Data & AI

    style Frontend fill:#3B82F6,color:#fff
    style Backend fill:#10B981,color:#fff
    style Data fill:#F59E0B,color:#fff
    style AI fill:#8B5CF6,color:#fff
    style Infra fill:#6B7280,color:#fff
Loading

View this diagram: open the .mmd file at docs/diagrams/architecture.mmd in Mermaid Live Editor

Component Breakdown

Component Responsibility
React SPA Chat interface, file management, admin dashboard
Nginx Reverse proxy, SPA serving, request routing to FastAPI
FastAPI REST API, SSE streaming, request validation, auth middleware
Auth Service JWT access tokens, refresh token rotation, OTP reset
RAG Pipeline Query embedding, vector search, prompt assembly, LLM call
Ingestion Service File parsing, text chunking, embedding, ChromaDB storage
Review Mode Full-document structured analysis via LLM
APScheduler Cron-based periodic re-indexing of all documents
PostgreSQL Users, files, chats, messages, access control, settings
ChromaDB Vector embeddings for semantic search
Ollama Self-hosted LLM inference and text embedding

Features

Document Upload and Indexing

Purpose: Users upload documents that are then processed into searchable vector embeddings.

User Workflow:

  1. User goes to My Files and uploads a file.
  2. The system validates the file type and size.
  3. A background task extracts text, chunks it, embeds each chunk, and stores the vectors in ChromaDB.
  4. The file status shows as "indexed" when complete.

Business Value: Documents become immediately queryable after upload. Users do not wait for manual processing.

Technical Summary: Magic-byte MIME validation (not just file extension), background async ingestion via FastAPI BackgroundTasks, sentence-based chunking via LlamaIndex SentenceSplitter, per-chunk Ollama embedding.


YouTube and URL Ingestion

Purpose: Ingest content from YouTube videos and web pages without downloading files manually.

User Workflow:

  1. User pastes a YouTube URL or a web URL.
  2. The system extracts captions (YouTube) or page text (web), then indexes it like any other document.

Business Value: Teams can quickly add research from videos and articles to their knowledge base.

Technical Summary: YouTube transcript extraction via yt-dlp with Whisper audio fallback. Web page content extraction via BeautifulSoup. Both routes feed into the same ingestion pipeline.


Q&A Chat Mode

Purpose: Users ask natural language questions and get answers grounded in selected documents.

User Workflow:

  1. User opens or creates a chat.
  2. Selects one or more indexed documents from the file panel.
  3. Asks a question.
  4. The system returns an answer with citations to the source document chunks.

Business Value: Replaces manual document search. Users get direct answers with source references.

Technical Summary: Vector similarity search against ChromaDB, cosine similarity threshold filtering, RAG prompt construction, SSE-based token streaming to the browser.


Document Review Mode

Purpose: Generates a structured analysis report for one or more documents.

User Workflow:

  1. User selects files and chooses Review mode.
  2. The system reads all indexed content and returns a structured report: summary, key sections, risk flags, and missing elements.

Business Value: Useful for contract review, compliance checks, and research synthesis.

Technical Summary: Aggregates all chunks for selected files, submits full document text to LLM with a structured analysis prompt, parses JSON response, renders as formatted markdown.


Streaming Chat Responses

Purpose: Chat answers stream to the screen token by token instead of waiting for the full response.

User Workflow: The user sends a message and sees the answer appear word by word, like a normal chat interface.

Business Value: Feels responsive. Users see progress and can act on partial answers.

Technical Summary: Server-Sent Events (SSE) over FastAPI StreamingResponse. Async generators yield tokens from Ollama, OpenAI, Anthropic, or Google streaming APIs.


Multi-Provider LLM Support

Purpose: Admins can switch between a local Ollama model and cloud API providers without redeploying.

User Workflow:

  1. Admin opens Settings.
  2. Selects provider (Ollama / Anthropic / OpenAI / Google).
  3. Enters API key for cloud providers.
  4. Saves. The new provider is active immediately.

Business Value: Teams can start with a free local model and switch to a better cloud model when needed.

Technical Summary: Settings stored in PostgreSQL. API keys encrypted with Fernet symmetric encryption before storage. Provider routing happens at runtime per request.


Admin User and Access Control

Purpose: Admins manage users, files, and who can access what.

User Workflow:

  1. Admin creates user accounts.
  2. Admin uploads files or approves user-uploaded files.
  3. Admin grants specific users access to specific files.
  4. Users only see files they own or have been granted access to.

Business Value: Documents stay compartmentalized. A user working on one project cannot read documents from another.

Technical Summary: Role-based access control (RBAC) with admin and user roles. File access tracked via a join table. Admin approval workflow for user-uploaded files.


Chat Sharing

Purpose: Users can share a read-only link to a specific chat conversation.

User Workflow: User enables sharing on a chat. A unique token is generated. Anyone with the link can read the conversation.

Business Value: Easy to share research findings externally without giving account access.

Technical Summary: Secure random token stored on the Chat record. Shared chats served via a public endpoint that requires no authentication.


Scheduled Re-indexing

Purpose: Admins can configure a cron schedule to automatically re-index all documents.

Technical Summary: APScheduler with an AsyncIOScheduler. Cron expression stored in the settings table. Scheduler reads the setting on startup and registers the job.


Technology Stack

Layer Technology Why Chosen
Frontend React 18 Component-based UI with hooks; good ecosystem
Styling Tailwind CSS Utility-first CSS; fast to build custom interfaces
Backend Python, FastAPI Async-native, fast, strong typing with Pydantic
Database PostgreSQL 15 Reliable, mature, JSONB support for message sources
Vector Store ChromaDB Simple self-hosted vector DB; no cloud dependency
Local AI Ollama Run open-source LLMs locally with zero cloud cost
Embedding nomic-embed-text (via Ollama) High-quality embeddings that run on CPU
Cloud LLM Anthropic / OpenAI / Google Optional cloud providers for better model quality
RAG Framework LlamaIndex (SentenceSplitter only) Used for sentence-aware document chunking
File Parsing pdfplumber, python-docx, ebooklib Reliable text extraction per format
Audio/Video faster-whisper, yt-dlp CPU-based transcription; YouTube audio extraction
Auth JWT + bcrypt Stateless access tokens with refresh token rotation
Encryption Fernet (cryptography library) Symmetric encryption for API keys at rest
Scheduling APScheduler Async-compatible cron scheduler
Rate Limiting SlowAPI Request rate limiting for auth endpoints
Logging Loguru Structured log rotation with compression
Migrations Alembic Version-controlled schema migrations
Deployment Docker Compose, Nginx Container-based; single-command deployment

Database Design

Entity Overview

erDiagram
    USERS ||--o{ FILES : uploads
    USERS ||--o{ REFRESH_TOKENS : has
    USERS ||--o{ PASSWORD_RESETS : has
    USERS ||--o{ FILE_ACCESS : granted
    USERS ||--o{ CHATS : creates
    USERS ||--o{ CHAT_ACCESS : granted

    FILES ||--o{ FILE_ACCESS : controls
    FILES ||--o{ CHAT_FILES : referenced_in

    CHATS ||--o{ MESSAGES : contains
    CHATS ||--o{ CHAT_ACCESS : controls
    CHATS ||--o{ CHAT_FILES : uses

    SETTINGS {
        key string
        value text
    }
Loading

View this diagram: open docs/diagrams/er-diagram.mmd in Mermaid Live Editor

Table Purposes

Table Purpose Key Relationships
users User accounts with roles and profile data Has files, chats, access grants, tokens
files Uploaded document records with indexing status Owned by user; accessed via file_access
file_access Explicit file permission grants Joins users and files
chats Chat sessions with optional share tokens Created by user; contains messages and files
chat_access Explicit chat permission grants Joins users and chats
messages Individual chat messages with role and sources Belongs to chat; stores source citations (JSONB)
chat_files Files linked to a chat session Joins chats and files
refresh_tokens Hashed refresh tokens with expiry Belongs to user
password_resets Hashed OTP codes for password recovery Belongs to user
settings Key-value store for runtime configuration No FK; global

User Roles and Permissions

Role Description Key Permissions
Admin System administrator Full access to all files and chats; user management; settings; file approval
User Standard team member Upload files; access granted files; create and manage their own chats

System Design Decisions

Design Patterns Used

  • Repository pattern (via SQLAlchemy async sessions): Database access is centralized through service functions, not scattered across route handlers.
  • Service layer: Business logic lives in services/, keeping route handlers thin and focused on HTTP concerns.
  • Background tasks: File ingestion and re-indexing run asynchronously via FastAPI's BackgroundTasks and APScheduler, keeping HTTP responses fast.
  • RBAC: Role-based access enforced at the dependency level. FastAPI Depends() injects role checks before route handlers run.
  • Refresh token rotation: Each token use issues a new token and invalidates the old one, limiting replay attack windows.

Multi-Provider LLM Approach

DocMind treats the LLM as a configurable service, not a hardcoded dependency. The active provider is read from the settings table on every request. This means admins can switch from Ollama to Anthropic without redeploying. API keys are encrypted at rest using Fernet symmetric encryption derived from the application secret key.

Security Approach

Authentication uses short-lived JWT access tokens (60 minutes) paired with long-lived refresh tokens stored as SHA-256 hashes in the database. Password hashing uses bcrypt. OTP codes for password reset are hashed before storage and expire after 15 minutes. File upload validation uses magic-byte detection, not just file extension checking. Rate limiting is applied to login and password-reset endpoints.

Scalability Approach

The application is stateless at the FastAPI layer. Horizontal scaling is possible by running multiple backend containers behind a load balancer, with shared PostgreSQL and ChromaDB. Ingestion and re-indexing run as background tasks, decoupled from the request lifecycle. ChromaDB persists on a Docker volume and can be replaced with a managed vector database for larger deployments.


API Overview

sequenceDiagram
    participant U as User
    participant FE as React Frontend
    participant API as FastAPI Backend
    participant SVC as Service Layer
    participant VEC as ChromaDB
    participant LLM as LLM (Ollama / Cloud)
    participant DB as PostgreSQL

    U->>FE: Sends chat message
    FE->>API: POST /chats/{id}/messages/stream
    API->>DB: Verify auth token, load chat
    API->>VEC: Embed query, similarity search
    VEC-->>API: Return matching chunks
    API->>SVC: Build RAG prompt from chunks
    SVC->>LLM: Stream prompt
    LLM-->>FE: SSE token stream
    API->>DB: Save complete message + sources
    FE-->>U: Renders answer with citations
Loading

View this diagram: open docs/diagrams/feature-flow.mmd in Mermaid Live Editor

Endpoint Categories

Category Purpose Auth Required
/auth Login, logout, register, token refresh, password reset Partial (login is public)
/files Upload, ingest, list, delete, reindex files Yes
/chats Create, read, send messages, share chats, stream Yes
/admin User management, file approval, access control Admin only
/settings Read and update runtime configuration Admin only
/shared Read-only access to shared chats via token No
/health Service health check No

Project Structure

docmind-github-public-version/
├── README.md                    <- You are here
├── ARCHITECTURE.md              <- Deep technical architecture
├── FEATURES.md                  <- Full feature documentation
├── DATABASE_OVERVIEW.md         <- Schema and data model
├── DEPLOYMENT_OVERVIEW.md       <- Infrastructure and deployment
├── SECURITY_OVERVIEW.md         <- Security model
├── CASE_STUDY.md                <- Client-facing project story
├── .env.example                 <- Environment variable template
├── src/
│   ├── backend/
│   │   ├── main.py              <- FastAPI app entry point
│   │   ├── config.py            <- Settings via pydantic-settings
│   │   ├── database.py          <- Async SQLAlchemy engine and session
│   │   ├── middleware.py        <- Error handling, request logging
│   │   ├── scheduler.py         <- APScheduler cron jobs
│   │   ├── seed.py              <- Admin user seed on startup
│   │   ├── schemas.py           <- Pydantic request/response schemas
│   │   ├── models/              <- SQLAlchemy ORM models
│   │   ├── routers/             <- FastAPI route handlers
│   │   ├── services/            <- Business logic and AI services
│   │   ├── alembic/             <- Database migration scripts
│   │   └── requirements.txt
│   └── frontend/
│       ├── src/
│       │   ├── App.jsx
│       │   ├── api/             <- Axios-based API client modules
│       │   ├── components/      <- Shared UI components
│       │   ├── pages/           <- Page-level components
│       │   └── context/         <- React context providers
│       ├── package.json
│       └── tailwind.config.js
├── docs/
│   ├── diagrams/
│   │   ├── architecture.mmd
│   │   ├── feature-flow.mmd
│   │   ├── tech-stack.mmd
│   │   └── er-diagram.mmd
│   └── screenshots/
│       └── README.md
└── LICENSE

Deployment Architecture

Developer Machine
      |
      | git push
      v
   VPS (Ubuntu)
      |
      | docker compose up -d
      v
+-------------------------------------+
|  Docker Compose Network             |
|                                     |
|  [Nginx + React]  --> [FastAPI]     |
|                         |           |
|                    [PostgreSQL]     |
|                    [ChromaDB]       |
|                    [Ollama]         |
+-------------------------------------+
      |
      | Nginx reverse proxy
      v
   Public Domain (HTTPS)
Environment Stack Purpose
Development Local Docker Compose Local dev with hot-reload
Production VPS + Docker Compose + Nginx Single-server production deploy

Development Journey

Key Challenges Solved

  • Audio and video ingestion without cloud dependency. The system needed to transcribe audio files locally. This required selecting a CPU-compatible Whisper model that balances speed and accuracy for typical document lengths, and integrating it into the same background ingestion pipeline used for text-based files.

  • Token streaming across four different LLM providers. Each provider has a different streaming API (Ollama uses NDJSON, Anthropic and OpenAI use chunked SSE, Google uses an async generator). The solution involved a unified streaming dispatcher that abstracts provider differences while preserving per-provider error handling.

  • Refresh token security without a cache layer. Maintaining token rotation security using only PostgreSQL required careful design: tokens are hashed before storage, rotation invalidates the old token atomically, and expired tokens are cleaned up on each new issuance.

  • File ingestion pipeline reliability. Ingestion involves text extraction, chunking, embedding (Ollama round-trip), and ChromaDB write. Any step can fail. The pipeline needed to handle partial failures gracefully, update file status accurately, and support admin-triggered re-indexing without data duplication.

  • MIME-based file validation. Relying on file extension for type detection is trivially bypassed. Using libmagic to read file headers catches disguised uploads early, before any parsing happens.

Architectural Decisions

ChromaDB vs. pgvector: Both were considered for vector storage. pgvector integrates into the existing PostgreSQL instance, which simplifies infrastructure. ChromaDB was chosen because its collection-based data model maps cleanly to per-user or per-project document sets, and it decouples vector scaling concerns from relational data concerns. The tradeoff is one more service to manage.

Single-server Docker Compose vs. Kubernetes: The target deployment is a single VPS for a small team. Kubernetes would add operational complexity without benefit at this scale. Docker Compose gives deterministic multi-service startup, health checks, and restart policies with minimal config. If the user base grows, the stateless FastAPI tier can be extracted and scaled independently.

Storing LLM provider config in the database vs. environment variables: Environment variables require a container restart to change. Storing provider config in PostgreSQL lets admins switch models or update API keys through a settings UI at runtime. The tradeoff is that the settings table becomes a dependency for every LLM call, mitigated by keeping setting reads lightweight.


Technical Skills Demonstrated

  • Backend: Async FastAPI with full request lifecycle management; SSE streaming; background task orchestration; Alembic migration workflow
  • AI/ML: RAG pipeline design (embed, retrieve, filter, prompt, generate); multi-provider LLM abstraction; local and cloud inference; audio transcription with Whisper
  • Database: PostgreSQL async access via SQLAlchemy; JSONB for structured message metadata; schema design for RBAC and access control
  • Security: JWT with refresh token rotation; bcrypt password hashing; Fernet encryption for secrets at rest; magic-byte file validation; rate limiting
  • Architecture: Service layer pattern; provider-agnostic LLM routing; background job scheduling; multi-format document ingestion pipeline
  • DevOps: Docker Compose orchestration with health checks; Nginx reverse proxy; multi-stage Docker builds; production-ready logging with rotation

Project Metrics

Metric Value
Total backend modules 17
API endpoints ~35
Database tables 10
Supported file types 7 (pdf, docx, epub, txt, mp3, wav, mp4) + YouTube + URL
LLM providers 4 (Ollama, Anthropic, OpenAI, Google)
Major features 8
Estimated complexity High

Screenshots

Live demo available on request. Screenshots provided during a private demo session.

Screen Description
Chat Interface Q&A with streaming responses and source citations
File Panel File list with indexing status indicators
Review Mode Structured document analysis with risk flags
Admin Dashboard User list, file management, and access control panel
Settings Panel LLM provider switcher and configuration

Case Study

See CASE_STUDY.md for the full project story including the challenge, approach, and outcomes.


Setup (Development Reference Only)

This public version contains skeleton code and is not runnable. Contact me for a private demo or to discuss the full implementation.

Prerequisites

  • Docker and Docker Compose
  • 4 GB RAM minimum (8 GB recommended for running Ollama locally)
  • Ubuntu 20.04+ or similar Linux host

Environment Variables

# Copy .env.example and fill in your values
DATABASE_URL=YOUR_VALUE_HERE
SECRET_KEY=YOUR_VALUE_HERE
OLLAMA_BASE_URL=YOUR_VALUE_HERE

License

This repository is published for portfolio and demonstration purposes only. All code is skeleton or placeholder and does not represent the production implementation. All proprietary business logic and algorithms are retained by the author.

2024 Abu Salah Mohammad Asif. All rights reserved.