AI Researcher — Red-Teaming, Model Evaluation & Adversarial Robustness
Bhimavaram, India · Portfolio · LinkedIn · Google Scholar · mpavangopinadh@gmail.com
AI researcher focused on red-teaming, model evaluation, and adversarial robustness, especially where AI models and systems fail. Works include jailbreaking, studying how models fail under representation shifts, RLHF failures, bias and fairness with accepted work at ICLR 2026 (AFAA Workshop) and ACL 2026 (EvalEval Workshop). Inaugural Adaption Research Grant grantee; B.Tech Computer Science 2026 (awaiting graduation). Experience building production AI and LLM systems (Applied AI Intern, Twimbit). Seeking full-time research roles; open to relocation.
Procedural Fairness Failures in RLHF from Preference Averaging ICLR 2026 — AFAA Workshop Identified a structural fairness failure where preference averaging in RLHF suppresses minority preference groups. Proposed Preference-Aware RLHF (PA-RLHF), improving alignment accuracy from 46.9% to 67.9% and reducing the fairness gap from 15.9 to 9.6 pp.
Are LLMs Safe Beyond Text? Do Emojis Expose Gaps in Safety Evaluation ACL 2026 — EvalEval Workshop Adversarial evaluation of Mistral-7B, Qwen-2-7B, Gemma-2-9B, and LLaMA-3-8B under emoji-augmented prompts. Found 0–10% jailbreak success rates, with statistically significant differences between models (χ² = 32.94, p < 0.001).
Regional Bias in Large Language Models AMRIT 2024 · arXiv:2601.16349 Designed FAZE, a context-neutral forced-choice framework for quantifying regional bias, applied across 10 LLMs (1,000 responses). Found 3.8× variation in bias scores that did not decrease with model scale.
Full publication list on Google Scholar.
Applied AI Intern, Twimbit (Singapore, remote) — Jul 2025 – Nov 2025 Built an end-to-end AI recruitment pipeline (n8n) that scores candidates against an HR-validated rubric and auto-schedules interviews, and a vendor research agent that turns daily enterprise/semiconductor/cloud news into structured executive briefings via LLM APIs.
Research: Model Evaluation, Adversarial testing, jailbreak benchmarking, model behaviour and patterns, RLHF failures, bias and fairness Stack: Python, Hugging Face, vLLM, OpenAI/Anthropic APIs, scipy/statsmodels, JSONL pipelines
- Official Grantee, Adaption Research Grant Program 2026 (Inaugural Cohort)
- Best Project Award, Dept. of Computer Science & Engineering, Vishnu Institute of Technology
- Participant, BlueDot Impact Technical AI Safety Course 2026
- 4× Hackathon Winner, 4× National-level Grand Finalist (incl. Amaravathi Quantum Valley Hackathon 2025, Smart India Hackathon 2024, 2× Google Hackathons)
- Selected, ACM India Summer & Winter Schools 2024 & 2025 (IIT Madras, IISc, IIT Gandhinagar, SRM AP)
mpavangopinadh@gmail.com · Portfolio · LinkedIn · Google Scholar · ORCID