Progressive delivery controller for ML model services: paired shadow analysis refuses a behaviorally broken model in 5.0 seconds with zero live exposure, staged canary ramp catches a +100ms regression at 5% traffic, and 10 of 10 identical-candidate rollouts promote with zero false rollbacks (all measured, results committed).
- New models ship on offline metrics and break behaviorally in production: here a candidate with a broken training feature (70.9% vs 83.5% holdout, looks plausible) is refused at the shadow stage because it flips 12% of the champion's high-confidence decisions, before any user sees it.
- Canary gates that alarm on noise get overridden and eventually turned off: every gate signal here breaches only on confidence bounds with hysteresis, and the committed benchmark shows 0 false rollbacks across 10 byte-identical rollouts.
- Rollback decisions nobody can explain afterwards: every evaluation writes the counts, bounds, and typed breach reasons to an append-only audit log, so "why did the gate kill my model" has an exact answer.
The expensive failure in ML deployment is not the model that crashes, it is the model that answers quickly, correctly formatted, and differently. A recommendation or risk model with a silently broken feature passes offline evaluation on the metrics its test set covers, then changes live decisions in ways the business discovers weeks later through complaint volume or loss rates. In the representative scenario built into this repo, a transaction-risk candidate trained with one dead feature keeps a plausible 70.9% holdout accuracy while flipping 12.4% of the cases the current model scores with 90%+ confidence, precisely the decisions users and downstream systems have learned to rely on.
model-canary-gate makes rollout safety a measured property. A gateway routes traffic between the champion and the candidate. In the mandatory shadow stage, live requests are mirrored to the candidate fire-and-forget (bounded, drop-counted backpressure; a mirror failure can never fail a user request) and every mirrored request yields a paired champion/candidate comparison. The gate admits the candidate to live traffic only if the Wilson lower bound on agreement clears a budget and high-confidence flips stay inside a stricter one. Then a staged canary ramp (5%, 25%, 50%) evaluates operational windows sized in candidate-arm requests: error rate by Wilson lower bound, prediction-rate shift by two-proportion confidence bound, latency by a dual rule requiring median AND tail regression. Consecutive breached windows, not single blips, trigger rollback, and the controller fails safe: a stage starved of traffic rolls back with evidence_timeout rather than promoting on no evidence.
Every claim is executed by committed benchmarks against real processes over real sockets. Ten rollouts of a byte-identical candidate: ten promotions, zero false rollbacks, median 15.1s to full promotion under the benchmark plan. The broken-feature candidate and a 50%-erroring candidate are both refused at shadow in 5.0 seconds with zero live exposure; a +100ms latency regression is rolled back at the 5% stage in 20.3 seconds. The gateway's safety tax, measured on 2 shared vCPUs: about 5.4ms added p50 for proxying, about 10.5ms with shadow mirroring active.
flowchart LR
U[clients] --> GW[gateway\nshadow mirror + weighted split]
GW -->|100% or 1-w| CH[champion server]
GW -->|mirrored or w| CA[candidate server]
GW --> ST[/paired agreement +\nper-arm windows/]
CTRL[rollout controller\nSHADOW to CANARY 5/25/50] -->|set mode| GW
ST -->|poll /stats| CTRL
CTRL -->|verdicts + evidence| AUD[(JSONL audit log)]
CTRL -->|breach confirmed| RB[rollback to champion]
Two design boundaries carry the guarantees: the gateway measures but never judges (hot path stays small and testable), and the controller judges but never serves (decisions are pure functions over stats snapshots, unit-testable without a network).
| Technology | Role in this project | Why chosen here |
|---|---|---|
| FastAPI + uvicorn | gateway and model servers | async mirroring needs first-class asyncio; the fire-and-forget shadow path is an asyncio.Task with bounded backpressure |
| httpx | gateway-to-backend and controller-to-gateway calls | one client API for production sockets and in-process ASGI transports, which is how the integration suite runs full rollouts with zero network flake |
| scikit-learn | champion and candidate variants | real models with a seeded data generator make disagreement genuine model behavior, reproducible from one seed |
| SQLite-free JSONL audit | decision log | append-only, greppable during an incident, no daemon; each entry carries the evidence for one verdict |
| pytest (+asyncio, +timeout) | 84 tests, 98% measured coverage | integration tests run the full controller/gateway/server topology in-process |
| Docker Compose | champion/candidate/gateway topology | the same processes the benchmarks measure, for interactive demos |
Prerequisites: Python 3.10+, git. Docker only for the compose demo.
git clone https://github.com/Panchalvedant13/model-canary-gate.git
cd model-canary-gate
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
canarygate train # seeded model zoo + holdout accuracies
# terminal 1: champion # terminal 2: candidate (pick a variant)
MODEL_NAME=champion MODEL_PORT=9001 python -m canarygate.server
MODEL_NAME=degraded MODEL_PORT=9002 python -m canarygate.server
# terminal 3: gateway # terminal 4: traffic + rollout
GATEWAY_PORT=9000 python -m canarygate.gateway
canarygate traffic --url http://127.0.0.1:9000 --rps 300 --duration 120 &
canarygate rollout --gateway http://127.0.0.1:9000 # exit 0 promoted, 3 rolled back
canarygate report # the audit trail
pytest # full suite, in-process rollouts included
python benchmark/overhead_bench.py # reproduce the overhead table
python benchmark/rollout_bench.py # reproduce the rollout matrix (~4 min)Or docker compose up -d --build after canarygate train for the same topology in containers (candidate variant via CANDIDATE_MODEL=degraded).
Methodology: benchmark/ scripts, real uvicorn processes over localhost sockets on 2 shared vCPUs (Intel Xeon class, 8 GB), 2,000 requests per cell, results committed in benchmark/results/. Shadow numbers include full mirroring of every request to the candidate.
xychart-beta
title "Client-observed p50 latency by path (ms)"
x-axis "concurrency" [4, 16, 32]
y-axis "p50 latency (ms)" 0 --> 180
line "direct to champion" [5.71, 20.06, 72.4]
line "via gateway" [11.11, 43.05, 93.0]
line "via gateway + shadow" [16.23, 73.2, 160.15]
| Path (c=4) | p50 | p95 | p99 | throughput |
|---|---|---|---|---|
| direct to champion | 5.71 ms | 9.54 ms | 13.09 ms | 615 rps |
| via gateway (proxied) | 11.11 ms | 18.35 ms | 21.50 ms | 332 rps |
| via gateway + shadow mirroring | 16.23 ms | 25.18 ms | 32.30 ms | 231 rps |
| Rollout scenario (committed run) | Outcome | Where | Time to decision |
|---|---|---|---|
| identical candidate, 10 trials | 10/10 PROMOTED | canary_50 | median 15.1s |
| retrained candidate, 3 trials | 3/3 PROMOTED | canary_50 | 10-15s |
| broken-feature candidate | ROLLED_BACK | shadow (zero live exposure) | 5.0s |
| 50% erroring candidate | ROLLED_BACK | shadow (zero live exposure) | 5.0s |
| +100ms latency candidate | ROLLED_BACK | canary_5 (5% exposure) | 20.3s |
Where it degrades and why: every process in these measurements (both model servers, the gateway, the load generator) shares 2 vCPUs, so the direct-path numbers degrade with concurrency from CPU contention and the shadow path pays double inference on the same cores; on separately provisioned serving instances the proxy hop cost (single-digit milliseconds) is the real overhead and the shadow tax lands on the candidate's hardware, not the client path. Time-to-decision scales linearly with traffic rate: at the benchmark's ~350-400 rps the 5% canary stage needs about 15 seconds to gather its 100 candidate-arm requests; at 10 rps it would need ten minutes, which is the correct behavior for an evidence-driven gate.
Two records in docs/adr/:
- ADR-001: shadow-first agreement gating; paired comparison on mirrored traffic buys per-request statistical power a canary cannot match, at the cost of doubled inference during the bounded shadow stage.
- ADR-002: the boring choice, fixed windows with Wilson bounds and hysteresis over always-valid sequential testing, defended with this repo's own false-rollback measurements and an explicit trigger for when sequential inference is the right tool instead.
- Tail-only latency regression detection. The latency gate requires median and tail to move together, because p95 alone on 100-request windows produced a measured 1.82x false positive between identical models. Trigger: windows of 1,000+ candidate requests, where tail quantiles stabilize enough to gate on separately.
- Statefulness and side effects in mirrored traffic. Shadow mirrors are fire-and-forget reads; candidates with write paths need request-level idempotency or a scrubbed mirror, which is a data-contract problem this gateway cannot solve generically.
- Multi-candidate tournaments and geo-split routing. One champion, one candidate, one decision. Trigger: the first real need to compare more than one candidate concurrently, which changes the statistics (multiplicity) as much as the routing.
- Auto-retraining or model registry integration. The gate decides safety, not lineage; wiring MLflow or a registry belongs to the deployment pipeline that invokes
canarygate rolloutand consumes its exit code.
No secrets exist locally (models are files, servers bind loopback); production deployments would inject backend URLs and any auth via environment or a secret manager, never the repo. The gateway logs decisions, counts, and latencies, never feature payloads: transaction features are exactly the kind of data that must not leak into log pipelines, so payload logging is structurally absent rather than toggled off. The audit log contains aggregates only. Containers run as a non-root user. The traffic generator sends synthetic, seeded feature vectors, so benchmarks never require real user data.
| Failure | Detection | Behavior | Recovery |
|---|---|---|---|
| Candidate server down or erroring during shadow | mirror failures counted per attempt | user requests unaffected (mirrors are fire-and-forget); gate refuses canary entry past the failure budget | fix candidate, rerun rollout |
| Candidate erroring during canary | per-window Wilson lower bound on 5xx | affected canary fraction sees 502s until confirm_breaches windows, then automatic rollback to champion | automatic |
| Traffic stops mid-rollout | stage deadline expires without evidence | fail safe: ROLLED_BACK with evidence_timeout; no promotion on no evidence |
rerun when traffic resumes |
| Candidate slower than mirror volume | pending-mirror set bounded, drops counted in /stats | memory stays bounded; sustained drops are visible evidence the candidate cannot absorb production load | treat drops as a shadow failure |
| Gateway crash | health endpoint, process supervisor | serving continues only if a supervisor restarts it; the gateway is deliberately stateless, mode is re-set by the controller | controller re-runs rollout; champion is default mode |
| Controller crash mid-rollout | gateway remains in last set mode | traffic keeps flowing in the last-known-safe configuration; audit log shows the last decision | operator inspects canarygate report, re-runs rollout |
The gate rolled back a byte-identical candidate. Twice, for two different reasons, and both rollbacks were the project's most valuable test failures.
The first came from the integration suite: positive_rate_shift_above_budget: |0.419 - 0.365| = 0.054 on roughly 60-request windows. The candidate was the champion's artifact byte for byte, so any detected "shift" was arithmetically guaranteed to be noise; the shift gate was comparing raw point estimates, violating the gate's own stated rule that rollbacks require confident evidence. The fix (commit 51e5e53) replaced the comparison with a two-proportion lower confidence bound, and the regression test replays the exact numbers from the failing run. The second came from the rollout benchmark on real sockets: p95_ratio=1.82 between identical models, pure queueing noise on a contended host, where 100-sample p95s diverge wildly while medians track within a few percent. The fix (commit 8d52980) requires median and tail to breach together, and tail-only detection moved to the documented out-of-scope list with an explicit trigger.
A third, different bug is preserved in commit b705d3f: the in-process integration harness hung because with in-memory ASGI transports every await completes synchronously, so the traffic-pump coroutine never reached a true suspension point and monopolized the event loop, starving the fire-and-forget mirror tasks and the controller. The diagnosis (asyncio only switches tasks at real suspension points) produced both a harness fix and a production improvement: the gateway's pending-mirror set is now bounded with drop accounting, so mirror backlog cannot grow without limit under overload.
- Sampled mirroring for shadow at high traffic (mirror 1-in-N with proportionally extended windows), activated by the ADR-001 cost trigger.
- Prometheus counters for gate evaluations, breaches, and mirror drops; the /stats endpoint already exposes the numbers, an exporter is mechanical.
- A
canarygate promote --manualpath that requires a typed reason, writing operator overrides into the same audit log as automatic decisions. - Weighted rollback memory: a candidate rolled back twice should need explicit human unblocking, not a third identical rollout.
- First metric to watch in production: the ratio of rollbacks at shadow versus canary stages; a rising canary share means behavioral coverage in shadow is decaying and agreement budgets need retuning.