I am an ML/LLM engineer and data scientist building reliable language-model systems—from post-training and cross-lingual representation analysis to RAG, information retrieval, and tool-calling agents.
My work emphasizes controlled experiments, exact grading, statistical uncertainty, failure analysis, and artifacts that others can rerun.
- Agent Reliability Lab — a deterministic, fault-injected benchmark for tool-calling agents with typed tools, policy-gated writes, replayable scenarios, and exact outcome grading.
- LatentSuff — a preregistered cross-lingual study of evidence-sufficiency signals in open-weight LLM representations across six languages and 29,206 paired QA constructions.
- LIMIT+ follow-up — an independent information-retrieval study separating BM25 candidate coverage from neural and symbolic constraint ranking.
- Reasoning Post-Training Lab — audit-first QLoRA SFT and DPO experiments with matched controls, locked selection, and hash-bound provenance.
Focus: ML/LLM systems · NLP and cross-lingual analysis · RAG, information retrieval and reranking · agent reliability · post-training
Tools: Python · PyTorch · Transformers · vLLM · CUDA · pytest · Docker

