| title | Vertex AI Endpoints (/invoke/*) vs. Cloud Run GPU |
|---|---|
| description | Architectural comparison of Google Cloud Vertex AI Dedicated Endpoints with arbitrary custom routes (invokeRoutePrefix="/*") vs. Serverless Cloud Run GPU for DiffusionGemma (dgemma). |
Reference for the two cloud runtimes behind dgem: why Vertex uses arbitrary custom routes, how the platforms
differ, and historical head-to-head receipts. For the current recommendation and step-by-step setup, follow
From Laptop to Production: Deploy on Cloud Run,
Production on Vertex AI and Gateway and routing.
Sections 2.1 and 5 are historical: they were measured on the retired L4 Vertex endpoint and an earlier serving image. Current measurements: Latency and capacity.
Standard Vertex AI Online Prediction (:predict and :rawPredict) binds a container deployment to a single fixed HTTP path (AIP_PREDICT_ROUTE). However, the dgemma container (structured_server.py on port 8080 fronting vLLM EngineCore on port 8000) exposes four distinct HTTP routes:
POST /v1/chat/completions— Structured 128-token diffusion decision envelope (answers+diagnostics).POST /v1/raw/chat/completions— Direct pass-through tovLLM's raw/v1/chat/completions.POST /v1/systemone— Multipart image + JSON decision route (JevBench/SystemOne).GET /health— Live container warmup phase, staged GiB telemetry, andvllm_readyflag.
By uploading the model with invokeRoutePrefix: "/*" and deploying it to a Vertex AI Dedicated Endpoint (dedicatedEndpointEnabled: true), Vertex AI forwards any non-root path under /invoke/<path> verbatim as /<path> to structured_server.py:
flowchart LR
subgraph Client["dgem CLI / dgemma-gateway / Web Studio"]
GW["dgemma-gateway\n(X-DGem-Backend: vertex)"]
end
subgraph Vertex["Vertex AI Dedicated Endpoint (*.prediction.vertexai.goog)"]
INV["/v1/projects/.../endpoints/{ID}/invoke/*"]
end
subgraph Container["dgemma Container (port 8080 -> vLLM 8000)"]
R1["POST /v1/chat/completions\n(Structured Decision Envelope)"]
R2["POST /v1/raw/chat/completions\n(Raw vLLM Pass-Through)"]
R3["POST /v1/systemone\n(SystemOne / JevBench)"]
R4["GET /health\n(Warmup & vLLM Readiness)"]
end
GW -->|"Bearer <OAuth2 cloud-platform>"| INV
INV -->|"/invoke/v1/chat/completions"| R1
INV -->|"/invoke/v1/raw/chat/completions"| R2
INV -->|"/invoke/v1/systemone"| R3
INV -->|"/invoke/health"| R4
| Dimension | Serverless Cloud Run GPU (dgemma) |
Vertex AI Dedicated Endpoint (dgemma-dedicated-g4) |
|---|---|---|
Scale-to-Zero (min=0) |
Yes (--min-instances=0) — Automatically scales down after 15 min idle ($0.00/hr when idle). |
No (minReplicaCount >= 1) — Bills per replica-hour while a model is deployed, until undeployed (make vertex-teardown). |
| Cold-Start / Provisioning | 2.3–2.7 min idle → warmed with Direct VPC egress (weights copied into memory in 64–83 s, overlapped with vLLM start-up). | ~10–15 min to deploy a model; replicas then stay warm (no wake-up). |
| Supported GPU Shapes | 1× NVIDIA L4 (24GB VRAM, 32Gi RAM) or 1× NVIDIA RTX Pro 6000 Blackwell (48GB VRAM, 80Gi RAM). |
g4-standard-48 + 1× RTX PRO 6000 (recommended; native NVFP4), plus L4 (g2-*), A100 (a2-*) and H100 (a3-*) machine types (the public images are tested only on RTX PRO 6000 and L4). |
| Multimodal (vision tower) | Single-GPU RTX PRO 6000 (80 GiB RAM) runs the model and the vision tower. | g4-standard-48 + RTX PRO 6000 runs NVFP4 weights and the vision tower on one GPU. |
| Routing Protocol | Direct HTTPS to https://dgemma-*.a.run.app/v1/chat/completions |
Arbitrary Custom Routes (invokeRoutePrefix: "/*") via https://<id>.<region>-<proj_num>.prediction.vertexai.goog/v1/projects/.../endpoints/<id>/invoke/v1/chat/completions |
| Authentication Token | OIDC Identity Token (gcloud auth print-identity-token, audience = Cloud Run URL) |
OAuth2 Access Token (gcloud auth print-access-token, scope = cloud-platform) |
| Payload Size Limit | 32 MiB (HTTP/1.1) / Unlimited streaming (HTTP/2) |
10–32 MiB on Dedicated Endpoints (*.prediction.vertexai.goog); bypasses the 1.5 MiB shared :predict limit |
| Health Probe Behavior | Polls GET /health (200 OK immediately so gateway can read live staging telemetry during scale-from-zero). |
Startup probe on GET /health (long timeout); readiness fields as on Cloud Run. |
| Enterprise MLOps Features | Revision traffic splitting, Direct Cloud Run IAP, Cloud Logging & Monitoring. | Model Registry versioning, Private Service Connect (PSC) endpoints, traffic-split canary rollouts (deployedModels/{id}/invoke/*), DCGM GPU AutoMetrics. |
We evaluated the exact same 30-case multi-domain decision suite (benchmarks/eval_dataset.jsonl across support, code_review, and security) on our live Vertex AI Dedicated Endpoint (<legacy-l4-endpoint>, g2-standard-16 · 1× NVIDIA L4 · /invoke/v1) and Serverless Cloud Run GPU (dgemma):
| Backend Target | Hardware Profile | Cold-Start / Wakeup | Single-Pass (N=1) Denoise |
4-Sample (N=4) Avg GPU Denoise |
Avg End-to-End Wall Time (30 cases) |
Proxy / Network Overhead | Multi-Domain Slot Accuracy | Receipt |
|---|---|---|---|---|---|---|---|---|
Vertex AI Dedicated Endpoint (/invoke/v1) |
g2-standard-16 (1× NVIDIA L4 24GB VRAM, 64GB RAM) |
0.0 s (minReplicaCount=1, always warm) |
195 ms (245 ms wall) |
490.0 ms (~122.5 ms/read) |
536.0 ms |
46.0 ms |
76.7% (23/30) |
benchmarks/results_vertex_l4_invoke.json |
Serverless Cloud Run GPU (dgemma) |
1× NVIDIA RTX Pro 6000 (48GB VRAM, 80Gi RAM) |
~121.8 s (0 → 1 scale-from-zero, $0.00/hr idle) |
171 ms (199 ms wall) |
427.3 ms (~106.8 ms/read) |
459.0 ms |
31.7 ms |
80.0% (24/30) |
benchmarks/results_cloudrun.json |
GCE VM (Raw vLLM without structured_server.py) |
g2-standard-8 (1× NVIDIA L4 24GB VRAM, 32GB RAM) |
0.0 s (dedicated VM) |
— | — | 1,968.7 ms (3.67× slower) |
— | 73.3% (22/30) |
benchmarks/results_gce_l4.json |
Local Apple Silicon Metal (diffgemma) |
Apple M-Series (q4 unified memory) |
0.0 s (local daemon) |
210 ms |
892.0 ms |
898.5 ms |
6.5 ms |
80.0% (24/30) |
benchmarks/results_local_metal_slot.json |
- Why
g2-standard-16(64 GBRAM) is Required for1× NVIDIA L4on Vertex AI:g2-standard-8(1× NVIDIA L4) provides only32 GBof host RAM. Staging the17.53 GiBdgemmasafetensors into/tmp/dgemma(tmpfsin RAM) whilevLLM/PyTorch allocates a17.53 GiBCPU load buffer requires~35.1 GiBof RAM, causing a container OOM kill ong2-standard-8.g2-standard-16attaches the exact same single1× NVIDIA L4GPU (24 GBVRAM) with64 GBof host RAM and16 vCPUs, eliminating the32 GBRAM bottleneck for negligible incremental CPU cost.
- 2-Second
/healthStartup for Vertex AI Health Probes:- Vertex AI Dedicated Endpoints fail
deploy-modelif port:8080/healthdoes not respond during container initialization.deploy/cloudrun/entrypoint.shstages only the<20 MBtokenizer +4 MBsafetensors headers synchronously (<2 seconds), launchesstructured_server.pyon:8080immediately so/healthreturns200 OK, and streams the17.53 GiBtensor bodies in the background across 64 HTTPS Range streams while_wait_for_upstream()holds early inference requests untilvLLMon:8000is ready.
- Vertex AI Dedicated Endpoints fail
Backend modes (vertex_first, vertex, cloudrun, local) and per-request overrides:
Gateway and routing. Deploying, swapping images and tearing down a dedicated endpoint:
Production on Vertex AI.
4. 4-Phase Head-to-Head Benchmark Matrix (historical: L4 endpoint, scripts/compare_vertex_vs_cloudrun.sh)
All empirical receipts are stored in benchmarks/results_head_to_head_vertex_vs_cloudrun.json, benchmarks/results_calibration_vertex_l4.json, and benchmarks/results_rerank_vertex_l4.json.
Concurrency (-w) |
Active vLLM Sequences (w × 4) |
Accuracy | GPU Denoise p50 |
GPU Denoise p90 |
GPU Denoise p99 |
Client Wall p50 |
Sustained Throughput | Operational Note |
|---|---|---|---|---|---|---|---|---|
w=1 (Sequential) |
4 sequences |
80.0% (24/30) |
499.4 ms |
517.1 ms |
576.0 ms |
534.0 ms |
1.87 dec/sec |
Lowest per-request latency (188 ms when N=1) |
w=4 (4 Workers) |
16 sequences (= MAX_NUM_SEQS) |
76.7% (23/30) |
717.4 ms |
769.8 ms |
827.6 ms |
786.0 ms |
5.44 dec/sec (30 cases in 5.51s) |
Sweet-spot concurrency for 1× NVIDIA L4 (24 GB VRAM) |
w=8 (8 Workers) |
32 sequences |
76.7% (23/30) |
1009.4 ms |
1025.7 ms |
1054.5 ms |
1047.0 ms |
7.49 dec/sec (30 cases in 4.01s) |
Peak throughput on a single L4 GPU replica (100% HTTP 200) |
w=16 (16 Workers) |
64 sequences |
— | — | — | — | — | — | Exceeds single-L4 24 GB KV-cache (MAX_NUM_SEQS=16); scale max-replica-count >= 2 or cap w <= 8 per L4 |
| Metric | Vertex AI Dedicated Endpoint (<legacy-l4-endpoint>) |
Serverless Cloud Run GPU (dgemma) |
|---|---|---|
Overall Suite Accuracy (50 cases) |
88.0% (44/50) |
82.0% (41/50) |
Chance-Corrected Accuracy (JevBench) |
83.05% |
74.57% |
10-Bin Expected Calibration Error (ECE) |
0.0470 (4.70%) |
0.0684 (6.84%) |
| Multi-Class Brier Score | 0.1944 |
0.2412 |
Composite JevBench v1.3.1 Score (GeoMean) |
73.37 (Intelligence: 84.10, Calibration: 84.85) |
71.18 |
Median Latency (p50) |
790.0 ms |
712.0 ms |
| Metric | Vertex AI Dedicated Endpoint (<legacy-l4-endpoint>) |
Serverless Cloud Run GPU (dgemma) |
|---|---|---|
| Simultaneous Slots per Forward Pass | 12 slots (10 passage grades + poisoned_passage + answer_present) |
12 slots (10 passage grades + poisoned_passage + answer_present) |
Mean Wall Latency (12 Slots / ~2,000 tokens) |
1,631.9 ms (~136 ms/slot) |
1,454.3 ms – 2,336.4 ms |
Continuous Expectation nDCG@10 |
0.8502 |
0.9265 |
FollowIR Policy-Flip p-MRR |
+0.8350 |
+0.7533 |
| Exact Tie Rate | 0.0% |
0.0% |
| RAG Prompt-Injection Poison Quarantine | 100.0% |
100.0% |