Skip to content
View Denina-V's full-sized avatar

Block or report Denina-V

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Denina-V/README.md

Denina Vincent

I build the infrastructure that sits between a promising model and a system
people can actually depend on — and then I measure whether it worked.

llm-gateway vitals-sentinel vecindex evalpower


The through-line

Anyone can get a model to produce an impressive demo. The interesting problems start immediately afterwards:

  • What happens when the provider 503s at 3am, and does your retry logic make it worse?
  • Who pays for this, and what stops one team's runaway loop from spending everyone's budget?
  • What faults get past your data checks — and how would you ever know?
  • Is that 6-point eval improvement real, or did you just ship noise?

Every repository below states a number, ships the code that reproduces it, and re-measures it in CI. Several of them say something I didn't expect and didn't want.


🛰 llm-gateway · platform

The unglamorous layer between an application and a very expensive API. Auth, per-key rate limits, hard spend caps, response caching, jittered retry, circuit breaker — behind one stable contract.

The design decision I'd defend in an interview: budgets reserve before the call and settle after. Fire 40 concurrent requests at a $0.50 key and 15 are admitted, 25 refused — all on held reservations, while spent_usd still reads 0.00. A post-hoc ledger sees that same zero and waves all forty through.

61 tests · 87% coverage · runs with no API key and no network

CI boots the built image and demands a 200 — which is how it caught a bug all 61 unit tests were blind to: the image shipped a config it never read, so every docker run request came back 401.

Python · FastAPI · Prometheus · Docker


🩺 vitals-sentinel · health data

A clinical data quality gate that has been scored, not just written. Most quality tools ask you to take their word for it. The interesting question isn't which rules they have — it's which faults get past them.

So it injects seven faults that actually happen on a hospital ward, records which rows each touched, and reports recall and precision per fault type: 100% recall on all seven, ≥99.4% precision.

The finding I'm proudest of is a mistake I made. My first drift baseline came from a different patient cohort — six drift findings where one existed. PSI was working perfectly; it just can't tell "the instrument changed" from "these are other people." Same-cohort reference, and one finding survives. CUSUM onset detection then located the change point within one hour and took precision from 11% → 99.4%.

49 tests · 93% coverage · no patient data ever reaches the model, pinned by a test

Python · pandas · scipy · Claude · data contracts


🔍 vecindex · systems

Brute force, IVF and HNSW written from scratch, then actually measured. Two results I did not expect:

HNSW lost. On 20k vectors it was beaten at every operating point — by IVF, and at high recall by plain brute force. Not a bug: IVF's inner loop is one BLAS matmul, HNSW's is Python-level pointer chasing. It does asymptotically less work in the slowest possible way. Growing the corpus to find the actual crossover put it at ~50,000 vectors — so on a 20k corpus, adding an HNSW index would have made the system almost twice as slow. Complexity class tells you where the curves cross, not whether you're past the crossing.

Keeping each node's M closest neighbours destroys the graph — 0.995 recall becomes 0.629, and raising search effort tenfold recovers nothing. Both graphs are fully connected with identical mean degree and no isolated nodes. Every health metric you'd think to check says the broken one is fine. Connectivity is not navigability.

61 tests · 97% coverage · the ablation runs in CI

Python · NumPy · HNSW · k-means


📉 evalpower · statistics

Your prompt eval cannot detect what you think it can.

Run both prompts, compare the numbers, ship the higher one. That procedure declares a winner 46.6% of the time when the two variants are identical. Detecting a genuine +5pp improvement at 80% power needs ~859 paired items — a 30-item eval finds it 3.5% of the time, against the 2.5% it fires on nothing at all. And peeking at the running result takes the false-positive rate from 4.5% to 36.3%.

Measured by simulating evals with a known true effect, thousands of times — which is the only way to check a test is right, since on real data you never know whether a detected difference was real.

The pairing result came out 1.1×–1.4×, not the dramatic win I expected. It's reported as measured.

59 tests · 97% coverage · CI re-runs all four experiments and asserts the headlines

Python · SciPy · McNemar · sequential testing


Domain adaptation for reproduction-rate prediction in an under-resourced health system, where the target label is missing in clusters rather than at random. EDA, a training pipeline, and a RAG layer over the supporting literature.

Python · scikit-learn · FAISS · Next.js


How I work

Numbers or it didn't happen. Every claim is reproducible from a clean checkout, and CI re-measures it on every push. If a change makes a gate blind to a fault class, the build goes red.

Report the result you got, not the one you wanted. HNSW losing and pairing being worth 1.2× are both in the READMEs, in bold, because a portfolio that only contains confirmations isn't evidence of measurement.

State the limits. Each README has a "what this deliberately is not" section — the in-process ledger that needs Redis before it scales out, the eval framework that doesn't model LLM-judge noise. Knowing where something breaks is part of having built it.

Tests that read like sentences. test_a_client_error_does_not_trip_the_breaker. test_no_patient_data_reaches_the_model. test_the_diversity_heuristic_beats_keeping_the_closest. If the name doesn't explain why the behaviour matters, the test isn't finished.


Currently interested in forward-deployed / platform engineering and health data roles.
Always happy to talk about circuit breakers, budget races, or why your eval can't see what you think it can.

Popular repositories Loading

  1. covid-uae-domain-adaptation covid-uae-domain-adaptation Public

    UAE COVID-19 EDA: missingness patterns, domain shift, and data scarcity — connecting OWID data to SSL/domain adaptation research

    Jupyter Notebook

  2. llm-gateway llm-gateway Public

    Production LLM gateway: per-key budgets enforced by reservation, circuit breaking, jittered retry, response caching and Prometheus metrics behind one stable contract. Runs with no API key.

    Python

  3. vitals-sentinel vitals-sentinel Public

    A clinical data-quality gate that has been scored, not just written: seven injected fault types, 100% recall measured on every CI run. Contract-driven checks, PSI drift with CUSUM onset detection, …

    Python

  4. Denina-V Denina-V Public

    Profile README

  5. evalpower evalpower Public

    Your prompt eval cannot detect what you think it can. Ship-the-higher-number declares a winner 46.6% of the time on identical variants; detecting +5pp at 80% power needs ~859 items. Measured by sim…

    Python

  6. vecindex vecindex Public

    Vector search from scratch: brute force, IVF and HNSW with an honest benchmark. HNSW lost at 20k vectors, and keeping each node's closest neighbours collapses recall 0.995 to 0.629 while every conn…

    Python