I design and build production LLM inference infrastructure on Kubernetes and NVIDIA GPUs. My current work covers vLLM-based model serving, traffic routing, performance and reliability, capacity planning, and multi-data-center resilience.
Previously, I built and scaled production backend and platform systems across SaaS, fintech, data privacy, and high-traffic consumer products.
- Production LLM serving with vLLM on NVIDIA GPUs
- Multi-model routing, quotas, fallbacks, and admission control
- Inference performance, benchmarking, and GPU capacity planning
- Kubernetes-based model deployment, observability, and reliability
- Multi-data-center inference architecture and failure recovery




