Skip to content
View shivamprasad1001's full-sized avatar
:shipit:
Currently fighting merge conflicts
:shipit:
Currently fighting merge conflicts

Block or report shivamprasad1001

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please donโ€™t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this userโ€™s behavior. Learn more about reporting abuse.

Report abuse
shivamprasad1001/README.md

Reinforcement Learning โ€ข Multi-Agent Systems โ€ข AI/ML Engineer

Typing SVG


About Me

I'm an AI/ML Engineer and final-year CS (AI & ML) student building intelligent systems at the intersection of Reinforcement Learning, Multi-Agent coordination, and applied AI engineering.

My current research focuses on cooperative multi-agent debate systems using PPO, human-in-the-loop orchestration, and scalable MARL architectures. I presented my work on adaptive multi-agent orchestration for automated code generation at NGAISL2026 (HRIT University).

Alongside research, I build production AI systems and automation tools through TriviLabs โ€” a web & AI agency serving Delhi NCR businesses.

Goal: Direct B.Tech โ†’ PhD in Reinforcement Learning & Multi-Agent Systems.



๐Ÿ“„ Research & Publications

Paper Venue Year
Adaptive Multi-Agent Orchestration with Human-in-the-Loop for Automated Code Generation NGAISL2026, HRIT University 2026

Current work: Cooperative Multi-Agent Debate System with PPO โ€” agents with opposing reward signals debate to convergence on complex reasoning tasks.


๐Ÿ”ฌ Research Focus

mindmap
  root((RL & MARL))
    Multi-Agent Systems
      Cooperative Debate with PPO
      MADDPG & QMIX
      MAPPO
      Emergent Communication
    Reinforcement Learning
      Policy Gradient Methods
      Sutton & Barto Foundations
      Human-in-the-Loop RL
      Reward Shaping
    Applied AI
      RAG & Retrieval Systems
      Agent Orchestration
      LLM Fine-tuning
      AI Automation Pipelines
    PhD Track
      University of Alberta target
      RLHF & Alignment
      Scalable MARL
Loading

๐Ÿš€ Signature Projects

๐Ÿค– Cooperative Multi-Agent Debate System

(In Progress)

Python PyTorch PPO

Multi-agent system where agents with opposing reward signals debate to convergence on complex tasks. Built on PPO with custom reward shaping for cooperative-competitive dynamics.

#reinforcement-learning #marl #ppo #debate

๐Ÿ“„ PaperMind AI

Stars

FastAPI LangChain React

Modern RAG-powered PDF chatbot โ€” upload documents, ask questions in natural language. Full-stack with FastAPI backend and React + Vite frontend.

#rag #langchain #pdf-chatbot #ai

๐Ÿ–ฅ๏ธ AI Desktop Assistant

Stars

Python NLP

Fully functional AI assistant with STT/TTS, intent detection, NER, and app control โ€” all using custom-trained models (no GPT wrapper).

#nlp #intent-classification #ner #desktop

๐ŸŽต MoodifyAI

Stars

OpenCV Flask

Detects emotions from facial expressions in real-time and recommends music tailored to the user's current mood.

#computer-vision #emotion-detection #music-recommendation

๐Ÿ” YOLO Custom Trainer

Stars

PyTorch YOLOv8

End-to-end object detection pipeline โ€” annotation with Label Studio, training, evaluation, and real-time Telegram bot deployment.

#yolov8 #computer-vision #object-detection

๐Ÿ› ๏ธ TriviLabs

Website

React n8n

Web & AI automation agency โ€” full-stack apps, AI workflows, and automation pipelines for Delhi NCR businesses.

#agency #ai-automation #web-development


๐Ÿ› ๏ธ Technical Stack

๐Ÿง  ML & Reinforcement Learning

Python PyTorch NumPy Pandas scikit-learn Hugging Face Gymnasium

๐Ÿ”ง AI Engineering & Agents

LangChain n8n FastAPI OpenCV RAG Groq

๐ŸŒ Full Stack & Cloud

React Node.js Supabase PostgreSQL Flask Linux


๐Ÿ“Š Stats


๐Ÿ—บ๏ธ 2026 Roadmap

gantt
    title Research & Career Roadmap 2026
    dateFormat YYYY-MM
    section RL Research
    Cooperative Multi-Agent Debate (PPO)   :active, 2026-01, 4M
    MARL paper submission                  :2026-05, 3M
    section PhD Preparation
    GRE 320+ target                        :active, 2026-01, 6M
    Professor outreach (~50 emails)        :2026-07, 3M
    PhD applications (18-22 programs)      :2026-09, 3M
    section Career
    AI Engineer role (Sarvam/Krutrim)      :2026-06, 6M
    TriviLabs โ‚น30-40k/month revenue       :active, 2026-01, 12M
Loading

๐Ÿค Connect

LinkedIn Portfolio TriviLabs HuggingFace Twitter Chat with Gwen


"Agents that debate are smarter than agents that agree."

Profile Views Followers

Pinned Loading

  1. gwen gwen Public

    Meet ๐ŸคGwen A RAG-powered personal AI. Ask about my research in RL & Multi-Agent Systems, TriviLabs, projects, and PhD journey. Built with FastAPI, React, Gemini, Groq & Pinecone.

    JavaScript

  2. papermind-ai papermind-ai Public

    PaperMind AI is a modern PDF document chat application powered by AI that allows you to upload PDF documents and have intelligent conversations about their content..

    TypeScript 5

  3. MoodifyAI MoodifyAI Public

    An AI-powered music recommendation system that detects emotions from facial expressions and suggests songs tailored to the user's mood.

    HTML 3 3

  4. emma.os emma.os Public

    Python