Skip to content

Latest commit

Β 

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Drishyamitra: Agentic Photos Evaluation and Segregation System

Node.js Python React Tailwind CSS Database Message Queue

Drishyamitra is an AI-powered photo management system designed to bring intelligence and automation to how users organize, search, and share their digital memories. Instead of manual sorting, Drishyamitra uses deep learning-based facial recognition, event-driven background queues, and natural language understanding to create an intuitive and efficient photo experience.


πŸ“– Table of Contents

  1. System Architecture
  2. Key Features
  3. Project Structure
  4. Prerequisites & Setup
  5. Core Workflows & Scenarios
  6. API Route Mapping
  7. Database Schema Details

πŸ—οΈ System Architecture

Drishyamitra is designed as a polyglot microservice system. Below is the architectural diagram showing component interfaces, data storage, and external integrations:

graph TD
    Client[React Client Vite] <-->|HTTP / Socket.io| Backend[Express API Server Node.js]
    Backend <-->|Session / Queue| Redis[(Redis / BullMQ)]
    Backend <-->|Metadata / Logs| Mongo[(MongoDB Atlas)]
    Backend -->|Internal HTTP| FaceService[Face Service Python / Flask]
    Backend -->|Media Store| Cloudinary[Cloudinary CDN]
    Backend -->|Delivery| Email[Nodemailer SMTP]
    Backend -->|Delivery| WA[whatsapp-web.js]
    Backend <-->|LLM Queries| Groq[Groq LPU API]
Loading

✨ Key Features

  • 🧠 Deep Learning Face Analysis: Integrates InsightFace (buffalo_l) to extract 512-dimension face vectors. Leverages an IoU-based Non-Maximum Suppression (NMS) threshold of 0.70 to eliminate duplicate detections.
  • 🏷️ Smart Face Labeling: Features a custom overlay canvas for manual annotations. Once a face is labeled, Drishyamitra calculates the average centroid embedding to automatically suggest and propagate matching labels on future uploads.
  • πŸ’¬ Conversational AI Chat: Powered by Groq LPU with multi-model routing (llama-3.1-8b-instant for quick searches and llama-3.3-70b-versatile for complex tool-calling workflows). Built-in rate limit fallbacks and regex parsing fail-safes (parseFailedGeneration) ensure continuous performance.
  • πŸ“¦ Intelligent Photo Delivery: Distributes photos via Gmail (Nodemailer) and WhatsApp (via whatsapp-web.js). If photo payloads exceed standard constraints (e.g. 25MB for Gmail, 100MB for WhatsApp), Drishyamitra dynamically prompts the user via Socket.io to compile them into an in-memory ZIP archive, uploads it to Cloudinary, and delivers a download proxy link.
  • 🧹 Auto-Cleanup Queue: Utilizes BullMQ background queues for asynchronous processing. A recurring worker detects expired ZIP files from Cloudinary every 24 hours and purges database delivery references to reclaim storage.

πŸ“ Project Structure

Drishyamitra/
β”œβ”€β”€ backend/                  # Node.js Express Server
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ agent/            # Conversational AI agent loop (agentLoop.js)
β”‚   β”‚   β”œβ”€β”€ config/           # Database, Redis, and Cloudinary configurations
β”‚   β”‚   β”œβ”€β”€ controllers/      # Route controllers (photo, face, chat, delivery)
β”‚   β”‚   β”œβ”€β”€ middlewares/      # Authentication & validator middlewares
β”‚   β”‚   β”œβ”€β”€ models/           # Mongoose Database Schemas
β”‚   β”‚   β”œβ”€β”€ routes/           # Express REST API routes
β”‚   β”‚   β”œβ”€β”€ services/         # Integrations (email, WhatsApp, ZIP compression)
β”‚   β”‚   └── workers/          # BullMQ queue processor and cleanup workers
β”‚   β”œβ”€β”€ package.json
β”‚   └── worker.js             # Main BullMQ entrypoint
β”‚
β”œβ”€β”€ face-service/             # Flask Python Microservice (InsightFace)
β”‚   β”œβ”€β”€ models/               # Cached ONNX models
β”‚   β”œβ”€β”€ routes/               # Health and /recognize routes
β”‚   β”œβ”€β”€ services/             # Bounding box & vector extraction algorithms
β”‚   β”œβ”€β”€ utils/                # Image decoding and URL streaming helpers
β”‚   β”œβ”€β”€ app.py                # Service entry point
β”‚   β”œβ”€β”€ config.py             # Port & ML configurations
β”‚   └── requirements.txt      # Scientific libraries (onnxruntime, insightface)
β”‚
β”œβ”€β”€ frontend/                 # React SPA (Vite + Tailwind CSS v4.0)
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ components/       # Shared UI parts (Sidebar, Dropzone, Modals)
β”‚   β”‚   β”œβ”€β”€ pages/            # Dashboard, Gallery, FaceLabeling, and Chat client
β”‚   β”‚   β”œβ”€β”€ router/           # Routing configuration
β”‚   β”‚   └── main.jsx
β”‚   β”œβ”€β”€ package.json
β”‚   └── vite.config.js
β”‚
β”œβ”€β”€ shared/                   # Shared types, constraints, and helper functions
└── docs/                     # API specification, feature maps, and architectures

βš™οΈ Prerequisites & Setup

Requirements

  • Node.js: v18.0.0+
  • Python: v3.9.0+ (Python v3.10 or v3.13 is recommended)
  • Redis: Local instance running on port 6379 (or Upstash instance URL)
  • MongoDB: Local MongoDB instance or MongoDB Atlas Connection URI

Environment Variables Configuration

Create a .env file in the root folder or configure individual .env files for the services:

Backend (backend/.env)

PORT=5000
MONGO_URI=mongodb://localhost:27017/drishyamitra
REDIS_URL=redis://127.0.0.1:6379

# JWT settings
JWT_SECRET=your_jwt_secret_key_here
JWT_EXPIRES_IN=7d

# Groq API Keys
GROQ_API_KEY=your_groq_api_key_here
GROQ_FAST_MODEL=llama-3.1-8b-instant
GROQ_REASONING_MODEL=llama-3.3-70b-versatile

# Cloudinary Integration
CLOUDINARY_CLOUD_NAME=your_cloudinary_cloud_name
CLOUDINARY_API_KEY=your_cloudinary_api_key
CLOUDINARY_API_SECRET=your_cloudinary_api_secret

# Delivery Services
GMAIL_USER=your_gmail_username@gmail.com
GMAIL_APP_PASS=your_gmail_app_password

# Face Microservice URL
FACE_SERVICE_URL=http://localhost:5001
WHATSAPP_SESSION_PATH=./whatsapp-session

Face Recognition Service (face-service/.env)

PORT=5001
FACE_MODEL=Facenet512
DETECTOR_BACKEND=retinaface
LOG_LEVEL=INFO
MIN_DETECTION_SCORE=0.40

Frontend (frontend/.env)

VITE_API_URL=http://localhost:5000
VITE_SOCKET_URL=http://localhost:5000

1. Face Recognition Service Setup (Python)

  1. Navigate to the face-service/ directory:
    cd face-service
  2. Create a Python virtual environment:
    python -m venv venv
  3. Activate the virtual environment:
    • PowerShell (Windows): .\venv\Scripts\Activate.ps1
    • CMD (Windows): .\venv\Scripts\activate.bat
    • Linux/macOS: source venv/bin/activate
  4. Install requirements:
    pip install -r requirements.txt
  5. Boot up the service:
    python app.py
    The face microservice will listen on http://localhost:5001.

2. Express Backend Setup (Node.js)

  1. Navigate to the backend/ directory:
    cd backend
  2. Install standard Node dependencies:
    npm install
  3. Start the Express API Server:
    npm run dev
    The backend will boot on http://localhost:5000.
  4. In a separate terminal, start the BullMQ worker processor:
    node worker.js

3. React Client Setup (React/Vite)

  1. Navigate to the frontend/ directory:
    cd frontend
  2. Install npm dependencies:
    npm install
  3. Run the development Vite compiler:
    npm run dev
    The React client will launch on http://localhost:5173.

πŸ”„ Core Workflows & Scenarios

πŸ“Έ Ingestion & Face Processing

  1. A user uploads photos through the frontend dropzone.
  2. The backend uploads the image binaries to Cloudinary and pushes an ingestion task to the Redis recognitionQueue.
  3. The BullMQ worker grabs the job, fetches the image URL, and makes an HTTP request to /recognize on the Face Service.
  4. The Face Service downloads the photo, detects bounding boxes via InsightFace, filters overlaps using NMS, computes embeddings, and returns them to the worker.
  5. The backend matches embeddings using Cosine Similarity against average Person centroids. It auto-labels faces exceeding the match threshold (0.60) and registers unknown faces.

πŸ’¬ Agentic Retrieval & Actions

  1. A user queries: "Email Mom her photos from our trip last year."
  2. The Express server routes the command to the Groq API (llama-3.3-70b-versatile).
  3. The model selects the necessary tools (e.g. searchPhotos, emailPhotos) and outputs JSON schema invocations.
  4. The backend evaluates the tools, runs photo checks, and returns structural items.
  5. If the total file attachment size is greater than 25MB, a Socket.io event delivery:zip-confirm is emitted to the React UI. Once confirmed, the system streams photos in parallel, compresses them to a ZIP file, uploads the archive to Cloudinary, and emails the download link.

πŸ”— API Route Mapping

Authentication Routes

  • POST /api/v1/auth/register - Create a new user account.
  • POST /api/v1/auth/login - Validate credentials and return JWT tokens.

Photo Management Routes

  • POST /api/v1/photos/upload - Upload image file and trigger recognition.
  • GET /api/v1/photos - Retrieve list of uploaded photos.
  • DELETE /api/v1/photos/:id - Delete photo from DB and Cloudinary.

Face Labeling Routes

  • GET /api/v1/faces/unlabeled - Get all unrecognized face crops.
  • POST /api/v1/faces/label - Assign a name to a specific face embedding.
  • GET /api/v1/faces/centroids - Retrieve average centroid profiles.

Conversational & Sharing Routes

  • POST /api/v1/chat - Chat with the agent loop.
  • POST /api/v1/delivery/share - Execute email or WhatsApp deliveries.
  • GET /api/v1/delivery/history - Fetch history logs of shares.
  • GET /api/v1/delivery/download/:deliveryId - Stream delivery ZIP file from Cloudinary.

πŸ—„οΈ Database Schema Details

Drishyamitra implements MongoDB modeling to maintain performance and data integrity:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           User                              β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ _id: ObjectId | email: String | passwordHash: String        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚ 1
                               β”‚
                               β”‚ *
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           Photo                             β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ _id: ObjectId | user: Ref(User) | cloudinaryUrl: String     β”‚
β”‚ bytes: Number | status: String ("processing", "completed")  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚ 1
                               β”‚
                               β”‚ *
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           Face                              β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ _id: ObjectId | photo: Ref(Photo) | bbox: Object            β”‚
β”‚ embedding: Array[512] | person: Ref(Person) | labeled: Boolβ”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚ *
                               β”‚
                               β”‚ 1
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           Person                            β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ _id: ObjectId | name: String | centroidEmbedding: Array[512]β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

βš–οΈ License

This project is licensed under the MIT License. See LICENSE for details.

About

AI-powered photo management system featuring automated facial recognition (InsightFace), conversational LLM assistants (Groq), and automated sharing (WhatsApp & Email).

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Contributors

Languages