I build the infrastructure that sits between a promising model and a system
people can actually depend on — and then I measure whether it worked.
Anyone can get a model to produce an impressive demo. The interesting problems start immediately afterwards:
- What happens when the provider 503s at 3am, and does your retry logic make it worse?
- Who pays for this, and what stops one team's runaway loop from spending everyone's budget?
- What faults get past your data checks — and how would you ever know?
- Is that 6-point eval improvement real, or did you just ship noise?
Every repository below states a number, ships the code that reproduces it, and re-measures it in CI. Several of them say something I didn't expect and didn't want.
🛰 llm-gateway · platform
The unglamorous layer between an application and a very expensive API. Auth, per-key rate limits, hard spend caps, response caching, jittered retry, circuit breaker — behind one stable contract.
The design decision I'd defend in an interview: budgets reserve before the call
and settle after. Fire 40 concurrent requests at a $0.50 key and 15 are
admitted, 25 refused — all on held reservations, while spent_usd still reads
0.00. A post-hoc ledger sees that same zero and waves all forty through.
61 tests · 87% coverage · runs with no API key and no network
CI boots the built image and demands a 200 — which is how it caught a bug all 61
unit tests were blind to: the image shipped a config it never read, so every
docker run request came back 401.
Python · FastAPI · Prometheus · Docker
🩺 vitals-sentinel · health data
A clinical data quality gate that has been scored, not just written. Most quality tools ask you to take their word for it. The interesting question isn't which rules they have — it's which faults get past them.
So it injects seven faults that actually happen on a hospital ward, records which rows each touched, and reports recall and precision per fault type: 100% recall on all seven, ≥99.4% precision.
The finding I'm proudest of is a mistake I made. My first drift baseline came from a different patient cohort — six drift findings where one existed. PSI was working perfectly; it just can't tell "the instrument changed" from "these are other people." Same-cohort reference, and one finding survives. CUSUM onset detection then located the change point within one hour and took precision from 11% → 99.4%.
49 tests · 93% coverage · no patient data ever reaches the model, pinned by a test
Python · pandas · scipy · Claude · data contracts
🔍 vecindex · systems
Brute force, IVF and HNSW written from scratch, then actually measured. Two results I did not expect:
HNSW lost. On 20k vectors it was beaten at every operating point — by IVF, and at high recall by plain brute force. Not a bug: IVF's inner loop is one BLAS matmul, HNSW's is Python-level pointer chasing. It does asymptotically less work in the slowest possible way. Growing the corpus to find the actual crossover put it at ~50,000 vectors — so on a 20k corpus, adding an HNSW index would have made the system almost twice as slow. Complexity class tells you where the curves cross, not whether you're past the crossing.
Keeping each node's M closest neighbours destroys the graph — 0.995 recall becomes 0.629, and raising search effort tenfold recovers nothing. Both graphs are fully connected with identical mean degree and no isolated nodes. Every health metric you'd think to check says the broken one is fine. Connectivity is not navigability.
61 tests · 97% coverage · the ablation runs in CI
Python · NumPy · HNSW · k-means
📉 evalpower · statistics
Your prompt eval cannot detect what you think it can.
Run both prompts, compare the numbers, ship the higher one. That procedure declares a winner 46.6% of the time when the two variants are identical. Detecting a genuine +5pp improvement at 80% power needs ~859 paired items — a 30-item eval finds it 3.5% of the time, against the 2.5% it fires on nothing at all. And peeking at the running result takes the false-positive rate from 4.5% to 36.3%.
Measured by simulating evals with a known true effect, thousands of times — which is the only way to check a test is right, since on real data you never know whether a detected difference was real.
The pairing result came out 1.1×–1.4×, not the dramatic win I expected. It's reported as measured.
59 tests · 97% coverage · CI re-runs all four experiments and asserts the headlines
Python · SciPy · McNemar · sequential testing
Domain adaptation for reproduction-rate prediction in an under-resourced health system, where the target label is missing in clusters rather than at random. EDA, a training pipeline, and a RAG layer over the supporting literature.
Python · scikit-learn · FAISS · Next.js
Numbers or it didn't happen. Every claim is reproducible from a clean checkout, and CI re-measures it on every push. If a change makes a gate blind to a fault class, the build goes red.
Report the result you got, not the one you wanted. HNSW losing and pairing being worth 1.2× are both in the READMEs, in bold, because a portfolio that only contains confirmations isn't evidence of measurement.
State the limits. Each README has a "what this deliberately is not" section — the in-process ledger that needs Redis before it scales out, the eval framework that doesn't model LLM-judge noise. Knowing where something breaks is part of having built it.
Tests that read like sentences. test_a_client_error_does_not_trip_the_breaker.
test_no_patient_data_reaches_the_model. test_the_diversity_heuristic_beats_keeping_the_closest.
If the name doesn't explain why the behaviour matters, the test isn't finished.
Currently interested in forward-deployed / platform engineering and health data roles.
Always happy to talk about circuit breakers, budget races, or why your eval can't see what you think it can.