Senior full-stack engineer building production AI systems end-to-end — from retrieval pipelines and model serving to the UI that people actually use. I care about latency, reliability, and observability.
- LLM applications — RAG, agents, tool-use, structured outputs
- Inference infra — serving open and closed models with sane cost/latency tradeoffs
- Eval & observability — because "vibes-based" prompting doesn't survive production
- Full-stack delivery — typed APIs, streaming UIs, and the boring glue that makes them ship


