| title | From Laptop to Production |
|---|---|
| description | The dgem developer journey: run locally, deploy a scale-to-zero Cloud Run GPU, promote to an always-warm Vertex AI endpoint, and operate it. Measured latency, cold start and capacity at each stage. |
This track takes dgem from your laptop to a production endpoint and keeps it running. Every number on these
pages is measured (sample sizes and receipts are linked); identifiers such as <PROJECT>, <REGION> and
<ENDPOINT_ID> are placeholders.
| Step | Page | What you get |
|---|---|---|
| 1 | Run on your laptop | A first decision in minutes on Apple Silicon or a local NVIDIA GPU |
| 2 | Use a remote GPU | Point the CLI, Studio and MCP at any hosted endpoint |
| 3 | Deploy on Cloud Run | Your own scale-to-zero GPU service: $0 when idle, ~2.5 min cold start |
| 4 | Production on Vertex AI | An always-warm endpoint with autoscaling replicas and zero cold start |
| 5 | Gateway and routing | One URL for users, automatic failover from Vertex to Cloud Run, IAP |
| 6 | Latency and capacity | Request design, engine settings and how to measure your own |
| 7 | Operations runbook | Health, promotion, rollback, teardown and troubleshooting |
| 8 | Monitoring and alerts | Scheduled health check, alert policies, what to do when they fire |
Hosting it yourself? Evaluate dgem on your own GPU covers GPU requirements, checking the install against known benchmark ranges, and a custom evaluation on your own labelled data.
| Stage | Platform | Use it for | Cost model | Cold start | Warm latency, 1 sample (p50) |
|---|---|---|---|---|---|
| Crawl | Apple Silicon (diffgemma, Metal, q4) or a local NVIDIA GPU |
Development, writing policies, offline and private data | Your machine | None | ~0.9 s on Metal (different engine; numbers don't transfer to GPU serving) |
| Walk | Cloud Run GPU, 1× RTX PRO 6000 | Batch evaluation, research, bursty internal use, failover | Per instance-second; $0 when idle | 2.3–2.7 min (lean image + Direct VPC egress) | 61 ms GPU / 153 ms end to end |
| Run | Vertex AI dedicated endpoint, g4-standard-48 + 1× RTX PRO 6000 |
Production: always warm, IAM, autoscaling, monitoring, images | Per replica-hour while deployed | None (min replicas ≥ 1) | 55 ms GPU / 150 ms end to end |
Latency: 3-question decision, 30 sequential requests, same-region client, serving image 504638d
(receipts). The same image, policies, CLI, HTTP
API and MCP server work on every stage.
- GPU: use a Blackwell GPU (RTX PRO 6000). The checkpoint is NVFP4 (4-bit); Blackwell runs FP4 natively, and the L4 was about 3× slower on every request type we measured. Other GPUs are untested (requirements).
- Production: a Vertex AI dedicated endpoint with minimum replicas ≥ 1, autoscaling on GPU duty cycle.
- Failover and batch: a scale-to-zero Cloud Run GPU service using the lean image and Direct VPC egress.
- One gateway (
dgem serve) in front of both, routingvertex_firstwith automatic failover. - Requests:
"samples": 1by default; opt in to4only when you need agreement/standard error; never"auto"on latency-sensitive paths. Put all questions about one input in one request. - Change safely: validate every new serving image next to production before switching traffic (runbook).
- Tools:
scripts/serving_speed.py(latency and concurrency),scripts/coldstart_probe.py(Cloud Run cold start),scripts/contract_diff.py(API contract). Failed requests are counted as errors, never as zero latency. - Single client in the same region: these are service-side characteristics, not end-user internet latencies.
- Receipts:
benchmarks/runs/20260927-image-parity(current image) andbenchmarks/runs/20260925-serving-speed(L4 comparison, earlier image).