I'm an AI/ML Engineer and final-year CS (AI & ML) student building intelligent systems at the intersection of Reinforcement Learning, Multi-Agent coordination, and applied AI engineering.
My current research focuses on cooperative multi-agent debate systems using PPO, human-in-the-loop orchestration, and scalable MARL architectures. I presented my work on adaptive multi-agent orchestration for automated code generation at NGAISL2026 (HRIT University).
Alongside research, I build production AI systems and automation tools through TriviLabs โ a web & AI agency serving Delhi NCR businesses.
Goal: Direct B.Tech โ PhD in Reinforcement Learning & Multi-Agent Systems.
| Paper | Venue | Year |
|---|---|---|
| Adaptive Multi-Agent Orchestration with Human-in-the-Loop for Automated Code Generation | NGAISL2026, HRIT University | 2026 |
Current work: Cooperative Multi-Agent Debate System with PPO โ agents with opposing reward signals debate to convergence on complex reasoning tasks.
mindmap
root((RL & MARL))
Multi-Agent Systems
Cooperative Debate with PPO
MADDPG & QMIX
MAPPO
Emergent Communication
Reinforcement Learning
Policy Gradient Methods
Sutton & Barto Foundations
Human-in-the-Loop RL
Reward Shaping
Applied AI
RAG & Retrieval Systems
Agent Orchestration
LLM Fine-tuning
AI Automation Pipelines
PhD Track
University of Alberta target
RLHF & Alignment
Scalable MARL
|
(In Progress) Multi-agent system where agents with opposing reward signals debate to convergence on complex tasks. Built on PPO with custom reward shaping for cooperative-competitive dynamics.
|
Modern RAG-powered PDF chatbot โ upload documents, ask questions in natural language. Full-stack with FastAPI backend and React + Vite frontend.
|
|
Fully functional AI assistant with STT/TTS, intent detection, NER, and app control โ all using custom-trained models (no GPT wrapper).
|
Detects emotions from facial expressions in real-time and recommends music tailored to the user's current mood.
|
|
End-to-end object detection pipeline โ annotation with Label Studio, training, evaluation, and real-time Telegram bot deployment.
|
Web & AI automation agency โ full-stack apps, AI workflows, and automation pipelines for Delhi NCR businesses.
|
gantt
title Research & Career Roadmap 2026
dateFormat YYYY-MM
section RL Research
Cooperative Multi-Agent Debate (PPO) :active, 2026-01, 4M
MARL paper submission :2026-05, 3M
section PhD Preparation
GRE 320+ target :active, 2026-01, 6M
Professor outreach (~50 emails) :2026-07, 3M
PhD applications (18-22 programs) :2026-09, 3M
section Career
AI Engineer role (Sarvam/Krutrim) :2026-06, 6M
TriviLabs โน30-40k/month revenue :active, 2026-01, 12M



