DocMind lets teams upload documents, web pages, and audio/video files, then ask questions about them using a local or cloud AI model. All data stays on your own server. No third-party cloud handles your documents unless you explicitly configure a cloud LLM provider.
Portfolio Note: This is a public showcase repository. It contains architecture documentation, design decisions, and skeleton code structure. The full production implementation is proprietary. The engineering depth, system design, and problem-solving approach demonstrated here reflect the real project.
Demo credentials — Username:
admin· Password:admin123
- Problem Statement
- Solution Overview
- Architecture Overview
- Features
- Technology Stack
- Database Design
- User Roles and Permissions
- System Design Decisions
- API Overview
- Project Structure
- Deployment Architecture
- Development Journey
- Technical Skills Demonstrated
- Project Metrics
- Screenshots
- Case Study
- Setup
- License
Teams working with sensitive documents — contracts, research papers, internal reports — cannot safely send those documents to public AI services like ChatGPT. But they still need to search, summarize, and ask questions about large volumes of content. Doing this manually is slow and error-prone, especially when the document library grows over time.
The problem is not just privacy. It is also fragmentation: files live in different places, different people have access to different documents, and there is no single place to query all of it.
DocMind is a self-hosted platform where an admin uploads and manages documents, controls who can see what, and provides a chat interface where users can ask questions that are answered by an AI using only the content of the selected documents. Nothing leaves the server unless the admin configures a cloud provider. The system supports local Ollama models out of the box, with optional cloud LLM providers (Anthropic, OpenAI, Google) configurable through a settings panel.
DocMind uses a three-tier architecture: a React frontend served by Nginx, a FastAPI backend handling all business logic and AI orchestration, and a data layer consisting of PostgreSQL for relational data and ChromaDB for vector storage. Ollama runs as a sidecar service for local embedding and inference.
flowchart TD
User["User / Admin\n(Browser)"]
subgraph Frontend["Frontend Layer"]
React["React 18\nSPA"]
Nginx["Nginx\nReverse Proxy"]
end
subgraph Backend["Backend Layer"]
FastAPI["FastAPI\nREST + SSE"]
Auth["Auth Service\nJWT + Refresh Tokens"]
RAG["RAG Pipeline\nRetrieval + Prompt Build"]
Ingest["Ingestion Service\nChunk + Embed"]
Review["Review Mode\nStructured Analysis"]
Sched["APScheduler\nCron Re-index"]
end
subgraph Data["Data Layer"]
PG["PostgreSQL\nUsers, Files, Chats"]
Chroma["ChromaDB\nVector Store"]
end
subgraph AI["AI Layer"]
Ollama["Ollama\nLocal LLM + Embeddings"]
Cloud["Cloud LLM\nAnthropic / OpenAI / Google"]
end
subgraph Infra["Infrastructure"]
Docker["Docker Compose\nContainer Orchestration"]
end
User --> Nginx --> React
React --> FastAPI
FastAPI --> Auth
FastAPI --> RAG
FastAPI --> Ingest
FastAPI --> Review
FastAPI --> Sched
RAG --> Chroma
RAG --> Ollama
RAG --> Cloud
Ingest --> Chroma
Ingest --> Ollama
FastAPI --> PG
Docker -.-> Frontend & Backend & Data & AI
style Frontend fill:#3B82F6,color:#fff
style Backend fill:#10B981,color:#fff
style Data fill:#F59E0B,color:#fff
style AI fill:#8B5CF6,color:#fff
style Infra fill:#6B7280,color:#fff
View this diagram: open the
.mmdfile at docs/diagrams/architecture.mmd in Mermaid Live Editor
| Component | Responsibility |
|---|---|
| React SPA | Chat interface, file management, admin dashboard |
| Nginx | Reverse proxy, SPA serving, request routing to FastAPI |
| FastAPI | REST API, SSE streaming, request validation, auth middleware |
| Auth Service | JWT access tokens, refresh token rotation, OTP reset |
| RAG Pipeline | Query embedding, vector search, prompt assembly, LLM call |
| Ingestion Service | File parsing, text chunking, embedding, ChromaDB storage |
| Review Mode | Full-document structured analysis via LLM |
| APScheduler | Cron-based periodic re-indexing of all documents |
| PostgreSQL | Users, files, chats, messages, access control, settings |
| ChromaDB | Vector embeddings for semantic search |
| Ollama | Self-hosted LLM inference and text embedding |
Purpose: Users upload documents that are then processed into searchable vector embeddings.
User Workflow:
- User goes to My Files and uploads a file.
- The system validates the file type and size.
- A background task extracts text, chunks it, embeds each chunk, and stores the vectors in ChromaDB.
- The file status shows as "indexed" when complete.
Business Value: Documents become immediately queryable after upload. Users do not wait for manual processing.
Technical Summary: Magic-byte MIME validation (not just file extension), background async ingestion via FastAPI BackgroundTasks, sentence-based chunking via LlamaIndex SentenceSplitter, per-chunk Ollama embedding.
Purpose: Ingest content from YouTube videos and web pages without downloading files manually.
User Workflow:
- User pastes a YouTube URL or a web URL.
- The system extracts captions (YouTube) or page text (web), then indexes it like any other document.
Business Value: Teams can quickly add research from videos and articles to their knowledge base.
Technical Summary: YouTube transcript extraction via yt-dlp with Whisper audio fallback. Web page content extraction via BeautifulSoup. Both routes feed into the same ingestion pipeline.
Purpose: Users ask natural language questions and get answers grounded in selected documents.
User Workflow:
- User opens or creates a chat.
- Selects one or more indexed documents from the file panel.
- Asks a question.
- The system returns an answer with citations to the source document chunks.
Business Value: Replaces manual document search. Users get direct answers with source references.
Technical Summary: Vector similarity search against ChromaDB, cosine similarity threshold filtering, RAG prompt construction, SSE-based token streaming to the browser.
Purpose: Generates a structured analysis report for one or more documents.
User Workflow:
- User selects files and chooses Review mode.
- The system reads all indexed content and returns a structured report: summary, key sections, risk flags, and missing elements.
Business Value: Useful for contract review, compliance checks, and research synthesis.
Technical Summary: Aggregates all chunks for selected files, submits full document text to LLM with a structured analysis prompt, parses JSON response, renders as formatted markdown.
Purpose: Chat answers stream to the screen token by token instead of waiting for the full response.
User Workflow: The user sends a message and sees the answer appear word by word, like a normal chat interface.
Business Value: Feels responsive. Users see progress and can act on partial answers.
Technical Summary: Server-Sent Events (SSE) over FastAPI StreamingResponse. Async generators yield tokens from Ollama, OpenAI, Anthropic, or Google streaming APIs.
Purpose: Admins can switch between a local Ollama model and cloud API providers without redeploying.
User Workflow:
- Admin opens Settings.
- Selects provider (Ollama / Anthropic / OpenAI / Google).
- Enters API key for cloud providers.
- Saves. The new provider is active immediately.
Business Value: Teams can start with a free local model and switch to a better cloud model when needed.
Technical Summary: Settings stored in PostgreSQL. API keys encrypted with Fernet symmetric encryption before storage. Provider routing happens at runtime per request.
Purpose: Admins manage users, files, and who can access what.
User Workflow:
- Admin creates user accounts.
- Admin uploads files or approves user-uploaded files.
- Admin grants specific users access to specific files.
- Users only see files they own or have been granted access to.
Business Value: Documents stay compartmentalized. A user working on one project cannot read documents from another.
Technical Summary: Role-based access control (RBAC) with admin and user roles. File access tracked via a join table. Admin approval workflow for user-uploaded files.
Purpose: Users can share a read-only link to a specific chat conversation.
User Workflow: User enables sharing on a chat. A unique token is generated. Anyone with the link can read the conversation.
Business Value: Easy to share research findings externally without giving account access.
Technical Summary: Secure random token stored on the Chat record. Shared chats served via a public endpoint that requires no authentication.
Purpose: Admins can configure a cron schedule to automatically re-index all documents.
Technical Summary: APScheduler with an AsyncIOScheduler. Cron expression stored in the settings table. Scheduler reads the setting on startup and registers the job.
| Layer | Technology | Why Chosen |
|---|---|---|
| Frontend | React 18 | Component-based UI with hooks; good ecosystem |
| Styling | Tailwind CSS | Utility-first CSS; fast to build custom interfaces |
| Backend | Python, FastAPI | Async-native, fast, strong typing with Pydantic |
| Database | PostgreSQL 15 | Reliable, mature, JSONB support for message sources |
| Vector Store | ChromaDB | Simple self-hosted vector DB; no cloud dependency |
| Local AI | Ollama | Run open-source LLMs locally with zero cloud cost |
| Embedding | nomic-embed-text (via Ollama) | High-quality embeddings that run on CPU |
| Cloud LLM | Anthropic / OpenAI / Google | Optional cloud providers for better model quality |
| RAG Framework | LlamaIndex (SentenceSplitter only) | Used for sentence-aware document chunking |
| File Parsing | pdfplumber, python-docx, ebooklib | Reliable text extraction per format |
| Audio/Video | faster-whisper, yt-dlp | CPU-based transcription; YouTube audio extraction |
| Auth | JWT + bcrypt | Stateless access tokens with refresh token rotation |
| Encryption | Fernet (cryptography library) | Symmetric encryption for API keys at rest |
| Scheduling | APScheduler | Async-compatible cron scheduler |
| Rate Limiting | SlowAPI | Request rate limiting for auth endpoints |
| Logging | Loguru | Structured log rotation with compression |
| Migrations | Alembic | Version-controlled schema migrations |
| Deployment | Docker Compose, Nginx | Container-based; single-command deployment |
erDiagram
USERS ||--o{ FILES : uploads
USERS ||--o{ REFRESH_TOKENS : has
USERS ||--o{ PASSWORD_RESETS : has
USERS ||--o{ FILE_ACCESS : granted
USERS ||--o{ CHATS : creates
USERS ||--o{ CHAT_ACCESS : granted
FILES ||--o{ FILE_ACCESS : controls
FILES ||--o{ CHAT_FILES : referenced_in
CHATS ||--o{ MESSAGES : contains
CHATS ||--o{ CHAT_ACCESS : controls
CHATS ||--o{ CHAT_FILES : uses
SETTINGS {
key string
value text
}
View this diagram: open docs/diagrams/er-diagram.mmd in Mermaid Live Editor
| Table | Purpose | Key Relationships |
|---|---|---|
| users | User accounts with roles and profile data | Has files, chats, access grants, tokens |
| files | Uploaded document records with indexing status | Owned by user; accessed via file_access |
| file_access | Explicit file permission grants | Joins users and files |
| chats | Chat sessions with optional share tokens | Created by user; contains messages and files |
| chat_access | Explicit chat permission grants | Joins users and chats |
| messages | Individual chat messages with role and sources | Belongs to chat; stores source citations (JSONB) |
| chat_files | Files linked to a chat session | Joins chats and files |
| refresh_tokens | Hashed refresh tokens with expiry | Belongs to user |
| password_resets | Hashed OTP codes for password recovery | Belongs to user |
| settings | Key-value store for runtime configuration | No FK; global |
| Role | Description | Key Permissions |
|---|---|---|
| Admin | System administrator | Full access to all files and chats; user management; settings; file approval |
| User | Standard team member | Upload files; access granted files; create and manage their own chats |
- Repository pattern (via SQLAlchemy async sessions): Database access is centralized through service functions, not scattered across route handlers.
- Service layer: Business logic lives in
services/, keeping route handlers thin and focused on HTTP concerns. - Background tasks: File ingestion and re-indexing run asynchronously via FastAPI's BackgroundTasks and APScheduler, keeping HTTP responses fast.
- RBAC: Role-based access enforced at the dependency level. FastAPI
Depends()injects role checks before route handlers run. - Refresh token rotation: Each token use issues a new token and invalidates the old one, limiting replay attack windows.
DocMind treats the LLM as a configurable service, not a hardcoded dependency. The active provider is read from the settings table on every request. This means admins can switch from Ollama to Anthropic without redeploying. API keys are encrypted at rest using Fernet symmetric encryption derived from the application secret key.
Authentication uses short-lived JWT access tokens (60 minutes) paired with long-lived refresh tokens stored as SHA-256 hashes in the database. Password hashing uses bcrypt. OTP codes for password reset are hashed before storage and expire after 15 minutes. File upload validation uses magic-byte detection, not just file extension checking. Rate limiting is applied to login and password-reset endpoints.
The application is stateless at the FastAPI layer. Horizontal scaling is possible by running multiple backend containers behind a load balancer, with shared PostgreSQL and ChromaDB. Ingestion and re-indexing run as background tasks, decoupled from the request lifecycle. ChromaDB persists on a Docker volume and can be replaced with a managed vector database for larger deployments.
sequenceDiagram
participant U as User
participant FE as React Frontend
participant API as FastAPI Backend
participant SVC as Service Layer
participant VEC as ChromaDB
participant LLM as LLM (Ollama / Cloud)
participant DB as PostgreSQL
U->>FE: Sends chat message
FE->>API: POST /chats/{id}/messages/stream
API->>DB: Verify auth token, load chat
API->>VEC: Embed query, similarity search
VEC-->>API: Return matching chunks
API->>SVC: Build RAG prompt from chunks
SVC->>LLM: Stream prompt
LLM-->>FE: SSE token stream
API->>DB: Save complete message + sources
FE-->>U: Renders answer with citations
View this diagram: open docs/diagrams/feature-flow.mmd in Mermaid Live Editor
| Category | Purpose | Auth Required |
|---|---|---|
| /auth | Login, logout, register, token refresh, password reset | Partial (login is public) |
| /files | Upload, ingest, list, delete, reindex files | Yes |
| /chats | Create, read, send messages, share chats, stream | Yes |
| /admin | User management, file approval, access control | Admin only |
| /settings | Read and update runtime configuration | Admin only |
| /shared | Read-only access to shared chats via token | No |
| /health | Service health check | No |
docmind-github-public-version/
├── README.md <- You are here
├── ARCHITECTURE.md <- Deep technical architecture
├── FEATURES.md <- Full feature documentation
├── DATABASE_OVERVIEW.md <- Schema and data model
├── DEPLOYMENT_OVERVIEW.md <- Infrastructure and deployment
├── SECURITY_OVERVIEW.md <- Security model
├── CASE_STUDY.md <- Client-facing project story
├── .env.example <- Environment variable template
├── src/
│ ├── backend/
│ │ ├── main.py <- FastAPI app entry point
│ │ ├── config.py <- Settings via pydantic-settings
│ │ ├── database.py <- Async SQLAlchemy engine and session
│ │ ├── middleware.py <- Error handling, request logging
│ │ ├── scheduler.py <- APScheduler cron jobs
│ │ ├── seed.py <- Admin user seed on startup
│ │ ├── schemas.py <- Pydantic request/response schemas
│ │ ├── models/ <- SQLAlchemy ORM models
│ │ ├── routers/ <- FastAPI route handlers
│ │ ├── services/ <- Business logic and AI services
│ │ ├── alembic/ <- Database migration scripts
│ │ └── requirements.txt
│ └── frontend/
│ ├── src/
│ │ ├── App.jsx
│ │ ├── api/ <- Axios-based API client modules
│ │ ├── components/ <- Shared UI components
│ │ ├── pages/ <- Page-level components
│ │ └── context/ <- React context providers
│ ├── package.json
│ └── tailwind.config.js
├── docs/
│ ├── diagrams/
│ │ ├── architecture.mmd
│ │ ├── feature-flow.mmd
│ │ ├── tech-stack.mmd
│ │ └── er-diagram.mmd
│ └── screenshots/
│ └── README.md
└── LICENSE
Developer Machine
|
| git push
v
VPS (Ubuntu)
|
| docker compose up -d
v
+-------------------------------------+
| Docker Compose Network |
| |
| [Nginx + React] --> [FastAPI] |
| | |
| [PostgreSQL] |
| [ChromaDB] |
| [Ollama] |
+-------------------------------------+
|
| Nginx reverse proxy
v
Public Domain (HTTPS)
| Environment | Stack | Purpose |
|---|---|---|
| Development | Local Docker Compose | Local dev with hot-reload |
| Production | VPS + Docker Compose + Nginx | Single-server production deploy |
-
Audio and video ingestion without cloud dependency. The system needed to transcribe audio files locally. This required selecting a CPU-compatible Whisper model that balances speed and accuracy for typical document lengths, and integrating it into the same background ingestion pipeline used for text-based files.
-
Token streaming across four different LLM providers. Each provider has a different streaming API (Ollama uses NDJSON, Anthropic and OpenAI use chunked SSE, Google uses an async generator). The solution involved a unified streaming dispatcher that abstracts provider differences while preserving per-provider error handling.
-
Refresh token security without a cache layer. Maintaining token rotation security using only PostgreSQL required careful design: tokens are hashed before storage, rotation invalidates the old token atomically, and expired tokens are cleaned up on each new issuance.
-
File ingestion pipeline reliability. Ingestion involves text extraction, chunking, embedding (Ollama round-trip), and ChromaDB write. Any step can fail. The pipeline needed to handle partial failures gracefully, update file status accurately, and support admin-triggered re-indexing without data duplication.
-
MIME-based file validation. Relying on file extension for type detection is trivially bypassed. Using libmagic to read file headers catches disguised uploads early, before any parsing happens.
ChromaDB vs. pgvector: Both were considered for vector storage. pgvector integrates into the existing PostgreSQL instance, which simplifies infrastructure. ChromaDB was chosen because its collection-based data model maps cleanly to per-user or per-project document sets, and it decouples vector scaling concerns from relational data concerns. The tradeoff is one more service to manage.
Single-server Docker Compose vs. Kubernetes: The target deployment is a single VPS for a small team. Kubernetes would add operational complexity without benefit at this scale. Docker Compose gives deterministic multi-service startup, health checks, and restart policies with minimal config. If the user base grows, the stateless FastAPI tier can be extracted and scaled independently.
Storing LLM provider config in the database vs. environment variables: Environment variables require a container restart to change. Storing provider config in PostgreSQL lets admins switch models or update API keys through a settings UI at runtime. The tradeoff is that the settings table becomes a dependency for every LLM call, mitigated by keeping setting reads lightweight.
- Backend: Async FastAPI with full request lifecycle management; SSE streaming; background task orchestration; Alembic migration workflow
- AI/ML: RAG pipeline design (embed, retrieve, filter, prompt, generate); multi-provider LLM abstraction; local and cloud inference; audio transcription with Whisper
- Database: PostgreSQL async access via SQLAlchemy; JSONB for structured message metadata; schema design for RBAC and access control
- Security: JWT with refresh token rotation; bcrypt password hashing; Fernet encryption for secrets at rest; magic-byte file validation; rate limiting
- Architecture: Service layer pattern; provider-agnostic LLM routing; background job scheduling; multi-format document ingestion pipeline
- DevOps: Docker Compose orchestration with health checks; Nginx reverse proxy; multi-stage Docker builds; production-ready logging with rotation
| Metric | Value |
|---|---|
| Total backend modules | 17 |
| API endpoints | ~35 |
| Database tables | 10 |
| Supported file types | 7 (pdf, docx, epub, txt, mp3, wav, mp4) + YouTube + URL |
| LLM providers | 4 (Ollama, Anthropic, OpenAI, Google) |
| Major features | 8 |
| Estimated complexity | High |
Live demo available on request. Screenshots provided during a private demo session.
| Screen | Description |
|---|---|
| Chat Interface | Q&A with streaming responses and source citations |
| File Panel | File list with indexing status indicators |
| Review Mode | Structured document analysis with risk flags |
| Admin Dashboard | User list, file management, and access control panel |
| Settings Panel | LLM provider switcher and configuration |
See CASE_STUDY.md for the full project story including the challenge, approach, and outcomes.
This public version contains skeleton code and is not runnable. Contact me for a private demo or to discuss the full implementation.
- Docker and Docker Compose
- 4 GB RAM minimum (8 GB recommended for running Ollama locally)
- Ubuntu 20.04+ or similar Linux host
# Copy .env.example and fill in your values
DATABASE_URL=YOUR_VALUE_HERE
SECRET_KEY=YOUR_VALUE_HERE
OLLAMA_BASE_URL=YOUR_VALUE_HEREThis repository is published for portfolio and demonstration purposes only. All code is skeleton or placeholder and does not represent the production implementation. All proprietary business logic and algorithms are retained by the author.
2024 Abu Salah Mohammad Asif. All rights reserved.