Skip to content

About

Progressive delivery controller for ML model services. Paired shadow analysis refuses behaviorally broken models in 5s with zero live exposure; staged 5/25/50% canary ramp with Wilson-bound gates and hysteresis rollback. Measured: 0 false rollbacks in 10 identical-model rollouts; +100ms regression caught at 5% traffic.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

model-canary-gate

Progressive delivery controller for ML model services: paired shadow analysis refuses a behaviorally broken model in 5.0 seconds with zero live exposure, staged canary ramp catches a +100ms regression at 5% traffic, and 10 of 10 identical-candidate rollouts promote with zero false rollbacks (all measured, results committed).

CI Coverage False rollbacks License

What this solves

  • New models ship on offline metrics and break behaviorally in production: here a candidate with a broken training feature (70.9% vs 83.5% holdout, looks plausible) is refused at the shadow stage because it flips 12% of the champion's high-confidence decisions, before any user sees it.
  • Canary gates that alarm on noise get overridden and eventually turned off: every gate signal here breaches only on confidence bounds with hysteresis, and the committed benchmark shows 0 false rollbacks across 10 byte-identical rollouts.
  • Rollback decisions nobody can explain afterwards: every evaluation writes the counts, bounds, and typed breach reasons to an append-only audit log, so "why did the gate kill my model" has an exact answer.

Executive summary

The expensive failure in ML deployment is not the model that crashes, it is the model that answers quickly, correctly formatted, and differently. A recommendation or risk model with a silently broken feature passes offline evaluation on the metrics its test set covers, then changes live decisions in ways the business discovers weeks later through complaint volume or loss rates. In the representative scenario built into this repo, a transaction-risk candidate trained with one dead feature keeps a plausible 70.9% holdout accuracy while flipping 12.4% of the cases the current model scores with 90%+ confidence, precisely the decisions users and downstream systems have learned to rely on.

model-canary-gate makes rollout safety a measured property. A gateway routes traffic between the champion and the candidate. In the mandatory shadow stage, live requests are mirrored to the candidate fire-and-forget (bounded, drop-counted backpressure; a mirror failure can never fail a user request) and every mirrored request yields a paired champion/candidate comparison. The gate admits the candidate to live traffic only if the Wilson lower bound on agreement clears a budget and high-confidence flips stay inside a stricter one. Then a staged canary ramp (5%, 25%, 50%) evaluates operational windows sized in candidate-arm requests: error rate by Wilson lower bound, prediction-rate shift by two-proportion confidence bound, latency by a dual rule requiring median AND tail regression. Consecutive breached windows, not single blips, trigger rollback, and the controller fails safe: a stage starved of traffic rolls back with evidence_timeout rather than promoting on no evidence.

Every claim is executed by committed benchmarks against real processes over real sockets. Ten rollouts of a byte-identical candidate: ten promotions, zero false rollbacks, median 15.1s to full promotion under the benchmark plan. The broken-feature candidate and a 50%-erroring candidate are both refused at shadow in 5.0 seconds with zero live exposure; a +100ms latency regression is rolled back at the 5% stage in 20.3 seconds. The gateway's safety tax, measured on 2 shared vCPUs: about 5.4ms added p50 for proxying, about 10.5ms with shadow mirroring active.

Architecture

flowchart LR
    U[clients] --> GW[gateway\nshadow mirror + weighted split]
    GW -->|100% or 1-w| CH[champion server]
    GW -->|mirrored or w| CA[candidate server]
    GW --> ST[/paired agreement +\nper-arm windows/]
    CTRL[rollout controller\nSHADOW to CANARY 5/25/50] -->|set mode| GW
    ST -->|poll /stats| CTRL
    CTRL -->|verdicts + evidence| AUD[(JSONL audit log)]
    CTRL -->|breach confirmed| RB[rollback to champion]
Loading

Two design boundaries carry the guarantees: the gateway measures but never judges (hot path stays small and testable), and the controller judges but never serves (decisions are pure functions over stats snapshots, unit-testable without a network).

Tech stack

Technology Role in this project Why chosen here
FastAPI + uvicorn gateway and model servers async mirroring needs first-class asyncio; the fire-and-forget shadow path is an asyncio.Task with bounded backpressure
httpx gateway-to-backend and controller-to-gateway calls one client API for production sockets and in-process ASGI transports, which is how the integration suite runs full rollouts with zero network flake
scikit-learn champion and candidate variants real models with a seeded data generator make disagreement genuine model behavior, reproducible from one seed
SQLite-free JSONL audit decision log append-only, greppable during an incident, no daemon; each entry carries the evidence for one verdict
pytest (+asyncio, +timeout) 84 tests, 98% measured coverage integration tests run the full controller/gateway/server topology in-process
Docker Compose champion/candidate/gateway topology the same processes the benchmarks measure, for interactive demos

Quickstart

Prerequisites: Python 3.10+, git. Docker only for the compose demo.

git clone https://github.com/Panchalvedant13/model-canary-gate.git
cd model-canary-gate
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"

canarygate train                                   # seeded model zoo + holdout accuracies

# terminal 1: champion        # terminal 2: candidate (pick a variant)
MODEL_NAME=champion MODEL_PORT=9001 python -m canarygate.server
MODEL_NAME=degraded MODEL_PORT=9002 python -m canarygate.server

# terminal 3: gateway         # terminal 4: traffic + rollout
GATEWAY_PORT=9000 python -m canarygate.gateway
canarygate traffic --url http://127.0.0.1:9000 --rps 300 --duration 120 &
canarygate rollout --gateway http://127.0.0.1:9000   # exit 0 promoted, 3 rolled back
canarygate report                                    # the audit trail

pytest                                   # full suite, in-process rollouts included
python benchmark/overhead_bench.py      # reproduce the overhead table
python benchmark/rollout_bench.py       # reproduce the rollout matrix (~4 min)

Or docker compose up -d --build after canarygate train for the same topology in containers (candidate variant via CANDIDATE_MODEL=degraded).

Performance under load

Methodology: benchmark/ scripts, real uvicorn processes over localhost sockets on 2 shared vCPUs (Intel Xeon class, 8 GB), 2,000 requests per cell, results committed in benchmark/results/. Shadow numbers include full mirroring of every request to the candidate.

xychart-beta
    title "Client-observed p50 latency by path (ms)"
    x-axis "concurrency" [4, 16, 32]
    y-axis "p50 latency (ms)" 0 --> 180
    line "direct to champion" [5.71, 20.06, 72.4]
    line "via gateway" [11.11, 43.05, 93.0]
    line "via gateway + shadow" [16.23, 73.2, 160.15]
Loading
Path (c=4) p50 p95 p99 throughput
direct to champion 5.71 ms 9.54 ms 13.09 ms 615 rps
via gateway (proxied) 11.11 ms 18.35 ms 21.50 ms 332 rps
via gateway + shadow mirroring 16.23 ms 25.18 ms 32.30 ms 231 rps
Rollout scenario (committed run) Outcome Where Time to decision
identical candidate, 10 trials 10/10 PROMOTED canary_50 median 15.1s
retrained candidate, 3 trials 3/3 PROMOTED canary_50 10-15s
broken-feature candidate ROLLED_BACK shadow (zero live exposure) 5.0s
50% erroring candidate ROLLED_BACK shadow (zero live exposure) 5.0s
+100ms latency candidate ROLLED_BACK canary_5 (5% exposure) 20.3s

Where it degrades and why: every process in these measurements (both model servers, the gateway, the load generator) shares 2 vCPUs, so the direct-path numbers degrade with concurrency from CPU contention and the shadow path pays double inference on the same cores; on separately provisioned serving instances the proxy hop cost (single-digit milliseconds) is the real overhead and the shadow tax lands on the candidate's hardware, not the client path. Time-to-decision scales linearly with traffic rate: at the benchmark's ~350-400 rps the 5% canary stage needs about 15 seconds to gather its 100 candidate-arm requests; at 10 rps it would need ten minutes, which is the correct behavior for an evidence-driven gate.

Architecture decisions

Two records in docs/adr/:

  • ADR-001: shadow-first agreement gating; paired comparison on mirrored traffic buys per-request statistical power a canary cannot match, at the cost of doubled inference during the bounded shadow stage.
  • ADR-002: the boring choice, fixed windows with Wilson bounds and hysteresis over always-valid sequential testing, defended with this repo's own false-rollback measurements and an explicit trigger for when sequential inference is the right tool instead.

Intentionally out of scope

  • Tail-only latency regression detection. The latency gate requires median and tail to move together, because p95 alone on 100-request windows produced a measured 1.82x false positive between identical models. Trigger: windows of 1,000+ candidate requests, where tail quantiles stabilize enough to gate on separately.
  • Statefulness and side effects in mirrored traffic. Shadow mirrors are fire-and-forget reads; candidates with write paths need request-level idempotency or a scrubbed mirror, which is a data-contract problem this gateway cannot solve generically.
  • Multi-candidate tournaments and geo-split routing. One champion, one candidate, one decision. Trigger: the first real need to compare more than one candidate concurrently, which changes the statistics (multiplicity) as much as the routing.
  • Auto-retraining or model registry integration. The gate decides safety, not lineage; wiring MLflow or a registry belongs to the deployment pipeline that invokes canarygate rollout and consumes its exit code.

Security and compliance

No secrets exist locally (models are files, servers bind loopback); production deployments would inject backend URLs and any auth via environment or a secret manager, never the repo. The gateway logs decisions, counts, and latencies, never feature payloads: transaction features are exactly the kind of data that must not leak into log pipelines, so payload logging is structurally absent rather than toggled off. The audit log contains aggregates only. Containers run as a non-root user. The traffic generator sends synthetic, seeded feature vectors, so benchmarks never require real user data.

Failure modes

Failure Detection Behavior Recovery
Candidate server down or erroring during shadow mirror failures counted per attempt user requests unaffected (mirrors are fire-and-forget); gate refuses canary entry past the failure budget fix candidate, rerun rollout
Candidate erroring during canary per-window Wilson lower bound on 5xx affected canary fraction sees 502s until confirm_breaches windows, then automatic rollback to champion automatic
Traffic stops mid-rollout stage deadline expires without evidence fail safe: ROLLED_BACK with evidence_timeout; no promotion on no evidence rerun when traffic resumes
Candidate slower than mirror volume pending-mirror set bounded, drops counted in /stats memory stays bounded; sustained drops are visible evidence the candidate cannot absorb production load treat drops as a shadow failure
Gateway crash health endpoint, process supervisor serving continues only if a supervisor restarts it; the gateway is deliberately stateless, mode is re-set by the controller controller re-runs rollout; champion is default mode
Controller crash mid-rollout gateway remains in last set mode traffic keeps flowing in the last-known-safe configuration; audit log shows the last decision operator inspects canarygate report, re-runs rollout

Hardest problem solved

The gate rolled back a byte-identical candidate. Twice, for two different reasons, and both rollbacks were the project's most valuable test failures.

The first came from the integration suite: positive_rate_shift_above_budget: |0.419 - 0.365| = 0.054 on roughly 60-request windows. The candidate was the champion's artifact byte for byte, so any detected "shift" was arithmetically guaranteed to be noise; the shift gate was comparing raw point estimates, violating the gate's own stated rule that rollbacks require confident evidence. The fix (commit 51e5e53) replaced the comparison with a two-proportion lower confidence bound, and the regression test replays the exact numbers from the failing run. The second came from the rollout benchmark on real sockets: p95_ratio=1.82 between identical models, pure queueing noise on a contended host, where 100-sample p95s diverge wildly while medians track within a few percent. The fix (commit 8d52980) requires median and tail to breach together, and tail-only detection moved to the documented out-of-scope list with an explicit trigger.

A third, different bug is preserved in commit b705d3f: the in-process integration harness hung because with in-memory ASGI transports every await completes synchronously, so the traffic-pump coroutine never reached a true suspension point and monopolized the event loop, starving the fire-and-forget mirror tasks and the controller. The diagnosis (asyncio only switches tasks at real suspension points) produced both a harness fix and a production improvement: the gateway's pending-mirror set is now bounded with drop accounting, so mirror backlog cannot grow without limit under overload.

Future work

  • Sampled mirroring for shadow at high traffic (mirror 1-in-N with proportionally extended windows), activated by the ADR-001 cost trigger.
  • Prometheus counters for gate evaluations, breaches, and mirror drops; the /stats endpoint already exposes the numbers, an exporter is mechanical.
  • A canarygate promote --manual path that requires a typed reason, writing operator overrides into the same audit log as automatic decisions.
  • Weighted rollback memory: a candidate rolled back twice should need explicit human unblocking, not a third identical rollout.
  • First metric to watch in production: the ratio of rollbacks at shadow versus canary stages; a rising canary share means behavioral coverage in shadow is decaying and agreement budgets need retuning.

About

Progressive delivery controller for ML model services. Paired shadow analysis refuses behaviorally broken models in 5s with zero live exposure; staged 5/25/50% canary ramp with Wilson-bound gates and hysteresis rollback. Measured: 0 false rollbacks in 10 identical-model rollouts; +100ms regression caught at 5% traffic.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages