Turns a factory's existing CCTV into a real-time workplace-safety monitor.
Five YOLO detectors, a FastAPI backend, a live operator dashboard, and Telegram alerting.
About this published copy. The inference layer is a stub:
backend/app/ml/inference.pyreturns no detections andcamera_worker.pyfeeds a blank frame instead of an RTSP stream. Trained weights and the production capture loop are not public. Everything else is the real code — API, database models, auth, alerting, the dashboard, the dataset and training pipelines, and the training runs undertraining-runs/. For a complete, runnable detector see helmet-detection-yolo.
Uzbek industrial plants already have hundreds of CCTV cameras installed, but the footage is only reviewed after an incident. One safety officer cannot watch forty feeds at once. Missing hard hats, phone use in restricted zones, falls and early-stage fires go unnoticed until they cost someone.
Sergak AI attaches to the RTSP streams of cameras that are already on site, runs YOLO detectors over them continuously, and pushes an annotated snapshot to a Telegram group within seconds of a violation — while writing every event to a database the plant can audit later.
Nothing leaves the plant's network: inference runs on-premise.
| Module | Detects | Status |
|---|---|---|
| helmet | Missing hard hat / PPE | Trained — 97.5% mAP@50, 82.2% mAP@50-95 |
| smoking | Smoking in prohibited areas | Trained — 87.1% mAP@50, 54.9% mAP@50-95 |
| phone | Phone use in restricted zones | Trained, benchmark rerun in progress |
| fall | Person falling / lying down | Trained, benchmark rerun in progress |
| fire_smoke | Early-stage fire and smoke | Trained, benchmark rerun in progress |
Reported figures come from Ultralytics model.val() runs; see Training for the setup behind them.
┌────────────────┐ RTSP ┌──────────────────────────┐
│ Existing CCTV │ ────────► │ camera_worker (per feed) │
│ Hikvision NVR │ │ reconnect · frame buffer │
└────────────────┘ └────────────┬─────────────┘
│
▼
┌──────────────────────────┐
│ inference (YOLOv8) │
│ helmet · phone · fall │
│ fire/smoke · smoking │
└────────────┬─────────────┘
│ detections
▼
┌──────────────────────────┐
│ alert_manager │
│ dedup · cooldown · route │
└────────────┬─────────────┘
┌───────────────────────┼────────────────────────┐
▼ ▼ ▼
┌──────────────────┐ ┌────────────────────┐ ┌───────────────────┐
│ events + media │ │ FastAPI REST API │ │ Telegram worker │
│ (database) │ │ → web dashboard │ │ (aiogram bot) │
└──────────────────┘ └────────────────────┘ └───────────────────┘
Monitoring
- Multi-camera RTSP ingest with automatic reconnect
- Hikvision NVR discovery — scans the subnet and registers channels
- Per-camera module assignment (which detectors run on which feed)
- Event deduplication and alert cooldown so one violation is not sent forty times
Operations
- Web dashboard: live view, event feed, analytics, floor plan, reports
- Departments and users — violations are routed to the responsible department
- Role-based access with JWT
- Google OAuth sign-in and OTP e-mail verification
- In-app chat between operators
Delivery
- Telegram bot: annotated snapshot + camera, module and timestamp
- E-mail notifications
- Docker Compose deployment
backend/
├── app/
│ ├── api/ auth · auth_google · cameras · chat · departments
│ │ discovery · events · modules · nvr · settings · users
│ ├── core/ auth · config · database · email · inference
│ │ otp · security · seed
│ ├── ml/ inference · camera_worker · alert_manager
│ ├── models/ SQLAlchemy models
│ └── workers/ telegram_bot.py
├── requirements.txt
├── config.yaml
└── Dockerfile
frontend/
├── index · login · register · cameras · events · analytics
│ modules · departments · users · reports · settings · chat · floorplan
└── assets/js/ api · app · auth · components · data · pages
kaska/ helmet dataset pipeline + training scripts
smoking/ smoking dataset pipeline + training scripts
docker-compose.yml
git clone https://github.com/Mardonaka05/sergak-ai.git
cd sergak-ai
cp backend/.env.example backend/.env # fill in the values below
docker compose up -dDashboard: http://localhost:8000 · API docs: http://localhost:8000/docs
Without Docker
cd backend
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
uvicorn app.main:app --reload| Variable | Description |
|---|---|
DATABASE_URL |
SQLite or MySQL connection string |
JWT_SECRET |
Random secret for token signing |
NVR_HOST / NVR_USER / NVR_PASS |
Hikvision NVR credentials for discovery |
RTSP_URLS |
Comma-separated stream URLs (if not using NVR discovery) |
BOT_TOKEN / TELEGRAM_CHAT_ID |
Telegram alerting |
SMTP_USER / SMTP_PASSWORD |
E-mail notifications and OTP |
GOOGLE_CLIENT_ID / GOOGLE_CLIENT_SECRET |
Google OAuth |
CONF_THRESHOLD |
Detection confidence cutoff |
DEVICE |
cuda:0 or cpu |
Model weights are not published. Put your own .pt files in backend/models_pt/, or train them with the pipelines in kaska/ and smoking/.
Datasets were assembled from public sources plus frames pulled from the deployment cameras themselves, labelled in CVAT (self-hosted via Docker) and exported in YOLO format.
python kaska/scripts/train.py --data kaska/merged/data.yaml \
--model yolov8n.pt \
--epochs 50 --imgsz 640 --batch 16Best helmet run — yolov8n, 640px, batch 16, 2 classes (helmet, no_helmet):
| Precision | Recall | mAP@50 | mAP@50-95 |
|---|---|---|---|
| 0.956 | 0.928 | 0.975 | 0.822 |
The single biggest gain came from adding frames sampled from the actual deployment cameras. Domain match beat dataset size and every hyperparameter change we tried.
Designed to run on-premise on an NVIDIA Jetson Orin NX alongside the plant's existing PoE camera network, so no video leaves the site.
- TensorRT export + INT8 quantization for Jetson
- Badge-based OCR worker identification for underground sites
- Rebuild the smoking dataset — current mAP@50-95 is too low to ship
- Per-shift reporting and export
MIT — see LICENSE.
Mardonbek Sulaymonqulov — AI / Computer Vision Engineer GitHub · mardonbeksulaymonqulov156@gmail.com