Drishyamitra is an AI-powered photo management system designed to bring intelligence and automation to how users organize, search, and share their digital memories. Instead of manual sorting, Drishyamitra uses deep learning-based facial recognition, event-driven background queues, and natural language understanding to create an intuitive and efficient photo experience.
- System Architecture
- Key Features
- Project Structure
- Prerequisites & Setup
- Core Workflows & Scenarios
- API Route Mapping
- Database Schema Details
Drishyamitra is designed as a polyglot microservice system. Below is the architectural diagram showing component interfaces, data storage, and external integrations:
graph TD
Client[React Client Vite] <-->|HTTP / Socket.io| Backend[Express API Server Node.js]
Backend <-->|Session / Queue| Redis[(Redis / BullMQ)]
Backend <-->|Metadata / Logs| Mongo[(MongoDB Atlas)]
Backend -->|Internal HTTP| FaceService[Face Service Python / Flask]
Backend -->|Media Store| Cloudinary[Cloudinary CDN]
Backend -->|Delivery| Email[Nodemailer SMTP]
Backend -->|Delivery| WA[whatsapp-web.js]
Backend <-->|LLM Queries| Groq[Groq LPU API]
- π§ Deep Learning Face Analysis: Integrates InsightFace (buffalo_l) to extract 512-dimension face vectors. Leverages an IoU-based Non-Maximum Suppression (NMS) threshold of
0.70to eliminate duplicate detections. - π·οΈ Smart Face Labeling: Features a custom overlay canvas for manual annotations. Once a face is labeled, Drishyamitra calculates the average centroid embedding to automatically suggest and propagate matching labels on future uploads.
- π¬ Conversational AI Chat: Powered by Groq LPU with multi-model routing (
llama-3.1-8b-instantfor quick searches andllama-3.3-70b-versatilefor complex tool-calling workflows). Built-in rate limit fallbacks and regex parsing fail-safes (parseFailedGeneration) ensure continuous performance. - π¦ Intelligent Photo Delivery: Distributes photos via Gmail (Nodemailer) and WhatsApp (via
whatsapp-web.js). If photo payloads exceed standard constraints (e.g. 25MB for Gmail, 100MB for WhatsApp), Drishyamitra dynamically prompts the user via Socket.io to compile them into an in-memory ZIP archive, uploads it to Cloudinary, and delivers a download proxy link. - π§Ή Auto-Cleanup Queue: Utilizes BullMQ background queues for asynchronous processing. A recurring worker detects expired ZIP files from Cloudinary every 24 hours and purges database delivery references to reclaim storage.
Drishyamitra/
βββ backend/ # Node.js Express Server
β βββ src/
β β βββ agent/ # Conversational AI agent loop (agentLoop.js)
β β βββ config/ # Database, Redis, and Cloudinary configurations
β β βββ controllers/ # Route controllers (photo, face, chat, delivery)
β β βββ middlewares/ # Authentication & validator middlewares
β β βββ models/ # Mongoose Database Schemas
β β βββ routes/ # Express REST API routes
β β βββ services/ # Integrations (email, WhatsApp, ZIP compression)
β β βββ workers/ # BullMQ queue processor and cleanup workers
β βββ package.json
β βββ worker.js # Main BullMQ entrypoint
β
βββ face-service/ # Flask Python Microservice (InsightFace)
β βββ models/ # Cached ONNX models
β βββ routes/ # Health and /recognize routes
β βββ services/ # Bounding box & vector extraction algorithms
β βββ utils/ # Image decoding and URL streaming helpers
β βββ app.py # Service entry point
β βββ config.py # Port & ML configurations
β βββ requirements.txt # Scientific libraries (onnxruntime, insightface)
β
βββ frontend/ # React SPA (Vite + Tailwind CSS v4.0)
β βββ src/
β β βββ components/ # Shared UI parts (Sidebar, Dropzone, Modals)
β β βββ pages/ # Dashboard, Gallery, FaceLabeling, and Chat client
β β βββ router/ # Routing configuration
β β βββ main.jsx
β βββ package.json
β βββ vite.config.js
β
βββ shared/ # Shared types, constraints, and helper functions
βββ docs/ # API specification, feature maps, and architectures
- Node.js:
v18.0.0+ - Python:
v3.9.0+(Pythonv3.10orv3.13is recommended) - Redis: Local instance running on port
6379(or Upstash instance URL) - MongoDB: Local MongoDB instance or MongoDB Atlas Connection URI
Create a .env file in the root folder or configure individual .env files for the services:
PORT=5000
MONGO_URI=mongodb://localhost:27017/drishyamitra
REDIS_URL=redis://127.0.0.1:6379
# JWT settings
JWT_SECRET=your_jwt_secret_key_here
JWT_EXPIRES_IN=7d
# Groq API Keys
GROQ_API_KEY=your_groq_api_key_here
GROQ_FAST_MODEL=llama-3.1-8b-instant
GROQ_REASONING_MODEL=llama-3.3-70b-versatile
# Cloudinary Integration
CLOUDINARY_CLOUD_NAME=your_cloudinary_cloud_name
CLOUDINARY_API_KEY=your_cloudinary_api_key
CLOUDINARY_API_SECRET=your_cloudinary_api_secret
# Delivery Services
GMAIL_USER=your_gmail_username@gmail.com
GMAIL_APP_PASS=your_gmail_app_password
# Face Microservice URL
FACE_SERVICE_URL=http://localhost:5001
WHATSAPP_SESSION_PATH=./whatsapp-sessionPORT=5001
FACE_MODEL=Facenet512
DETECTOR_BACKEND=retinaface
LOG_LEVEL=INFO
MIN_DETECTION_SCORE=0.40VITE_API_URL=http://localhost:5000
VITE_SOCKET_URL=http://localhost:5000- Navigate to the
face-service/directory:cd face-service - Create a Python virtual environment:
python -m venv venv
- Activate the virtual environment:
- PowerShell (Windows):
.\venv\Scripts\Activate.ps1 - CMD (Windows):
.\venv\Scripts\activate.bat - Linux/macOS:
source venv/bin/activate
- PowerShell (Windows):
- Install requirements:
pip install -r requirements.txt
- Boot up the service:
The face microservice will listen on
python app.py
http://localhost:5001.
- Navigate to the
backend/directory:cd backend - Install standard Node dependencies:
npm install
- Start the Express API Server:
The backend will boot on
npm run dev
http://localhost:5000. - In a separate terminal, start the BullMQ worker processor:
node worker.js
- Navigate to the
frontend/directory:cd frontend - Install npm dependencies:
npm install
- Run the development Vite compiler:
The React client will launch on
npm run dev
http://localhost:5173.
- A user uploads photos through the frontend dropzone.
- The backend uploads the image binaries to Cloudinary and pushes an ingestion task to the Redis
recognitionQueue. - The BullMQ worker grabs the job, fetches the image URL, and makes an HTTP request to
/recognizeon the Face Service. - The Face Service downloads the photo, detects bounding boxes via InsightFace, filters overlaps using NMS, computes embeddings, and returns them to the worker.
- The backend matches embeddings using Cosine Similarity against average Person centroids. It auto-labels faces exceeding the match threshold (
0.60) and registers unknown faces.
- A user queries: "Email Mom her photos from our trip last year."
- The Express server routes the command to the Groq API (
llama-3.3-70b-versatile). - The model selects the necessary tools (e.g.
searchPhotos,emailPhotos) and outputs JSON schema invocations. - The backend evaluates the tools, runs photo checks, and returns structural items.
- If the total file attachment size is greater than 25MB, a Socket.io event
delivery:zip-confirmis emitted to the React UI. Once confirmed, the system streams photos in parallel, compresses them to a ZIP file, uploads the archive to Cloudinary, and emails the download link.
POST /api/v1/auth/register- Create a new user account.POST /api/v1/auth/login- Validate credentials and return JWT tokens.
POST /api/v1/photos/upload- Upload image file and trigger recognition.GET /api/v1/photos- Retrieve list of uploaded photos.DELETE /api/v1/photos/:id- Delete photo from DB and Cloudinary.
GET /api/v1/faces/unlabeled- Get all unrecognized face crops.POST /api/v1/faces/label- Assign a name to a specific face embedding.GET /api/v1/faces/centroids- Retrieve average centroid profiles.
POST /api/v1/chat- Chat with the agent loop.POST /api/v1/delivery/share- Execute email or WhatsApp deliveries.GET /api/v1/delivery/history- Fetch history logs of shares.GET /api/v1/delivery/download/:deliveryId- Stream delivery ZIP file from Cloudinary.
Drishyamitra implements MongoDB modeling to maintain performance and data integrity:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β User β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β _id: ObjectId | email: String | passwordHash: String β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β 1
β
β *
ββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββ
β Photo β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β _id: ObjectId | user: Ref(User) | cloudinaryUrl: String β
β bytes: Number | status: String ("processing", "completed") β
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β 1
β
β *
ββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββ
β Face β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β _id: ObjectId | photo: Ref(Photo) | bbox: Object β
β embedding: Array[512] | person: Ref(Person) | labeled: Boolβ
ββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β *
β
β 1
ββββββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββ
β Person β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β _id: ObjectId | name: String | centroidEmbedding: Array[512]β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
This project is licensed under the MIT License. See LICENSE for details.