Computer Vision Engineer | Machine Learning Researcher
I build intelligent vision systems and the machine learning models that power them. Passionate about bringing AI out of the lab and into the real world, I specialize in solving complex problems across Computer Vision, Edge AI, and Generative AI.
π Check out my full interactive portfolio: akshaysatyam2.github.io/akshaysatyam2
- Computer Vision: Edge AI, Object Detection & Tracking, 3D Vision, Generative Models, Biometrics, Real-time Analysis.
- Machine Learning & GenAI: LLMs, Retrieval-Augmented Generation (RAG), Deep Learning, MLOps, NLP.
A lightweight, fully native alternative to OBS Studio that eliminates the need for physical green screens. It uses WebRTC for capture and pipes it into a custom async Python backend, utilizing YOLO26 Nano Segmentation and ONNX Runtime (CUDA/OpenVINO) for real-time, hardware-accelerated alpha compositing at 30-60 FPS.
A comprehensive journey through 3D vision concepts, moving from epipolar geometry and PointNet architectures to real-time LiDAR point cloud processing. Includes a production-ready pipeline using YOLO26 to perform 2D detection and project it into 3D space with depth estimation for real-time 4K video tracking.
A real-time, edge-optimized AI system for pet monitoring and behavior analysis. Engineered to run flawlessly on resource-constrained hardware.
A Denoising Diffusion Probabilistic Model (DDPM) built entirely from scratch to deeply understand the mathematics behind modern AI image generation.
High-accuracy object detection pipelines for real-world environments, tackling everything from smart city traffic violations to robust drone tracking using YOLO and SSD models.
A robust QR code detection and decoding pipeline using OpenCV, designed to handle challenging real-world conditions including noisy, rotated, or low-contrast images.
A smart surveillance system leveraging YOLO and MediaPipe pose estimation to track head orientation and detect suspicious behavior during automated exam proctoring.
Facial recognition and OCR engines designed from the ground up, focusing on custom neural architectures rather than high-level abstractions.
A production-grade Retrieval-Augmented Generation (RAG) system combining dense/sparse hybrid search (Qdrant + BM25), adaptive document-scale chunking, hierarchical section breadcrumbs, knowledge graph expansion, and cross-encoder re-ranking (ms-marco-MiniLM-L-6-v2) to deliver grounded, hallucination-free answers from technical books, research papers, and complex documents.
A conversational AI application powered by a fine-tuned GPT-2 model to clone personal chatting styles, complete with an interactive Streamlit interface.
Looking to collaborate on cutting-edge vision systems or machine learning projects? Let's talk.
- Email: akshaysatyam2003@gmail.com
- LinkedIn: akshaysatyam2



.png)
.png)
.png)
.png)




