| title | Step 5: Gateway and Routing |
|---|---|
| description | Put one dgem gateway (dgem serve) in front of Vertex AI and Cloud Run: backend modes, automatic failover, per-request overrides on every surface, IAP access and deployment. |
The gateway (dgem serve) is what users and agents talk to. It serves Decision Studio, the HTTP API
(/api/decide, /v1/systemone, /v1/chat/completions), the MCP server (/mcp) and traces, and routes each
request to a GPU backend. Locally it runs on http://localhost:8090; in the cloud it runs as the Cloud Run service
dgemma-gateway behind IAP (placeholder https://<your-dgem-gateway>).
This is the reference table for backend selection; other pages link here.
| Mode | Sends requests to | When the primary is unavailable | Wake-up | Use it for |
|---|---|---|---|---|
vertex_first (default) |
Vertex AI dedicated endpoint | Fails over to Cloud Run (endpoint updating, scaled to 0, or erroring) | None while Vertex is warm | Studio, MCP agents, production APIs |
vertex |
Vertex AI only | Returns the error | None | Strict latency SLAs; testing the endpoint itself |
cloudrun |
Cloud Run GPU only | Holds and retries while the service wakes (--wakeup-timeout, default 10 min) |
~2.5 min after idle | Batch jobs, experiments, $0-idle use |
local |
Local diffgemma (--local) |
Returns the error | None | Laptop development |
Warm latencies are in Production on Vertex AI and
Deploy on Cloud Run. GET /api/backend-config lists the backends the gateway is
configured for (available_backends); the Studio hides the backend switcher when only one is configured.
| Setting | Default | Effect |
|---|---|---|
--backends vertex_first,cloudrun,local (DGEM_BACKENDS) |
derived from the configured URLs | Explicit allow-list. Startup fails if a listed backend lacks its URL or --default-backend is not in the list. |
--enable-admin-api (DGEM_ADMIN_API=1) |
off | Allows POST /api/backend-config (gateway-wide default backend, Vertex URL, local URL) and POST /api/vertex/deploy / teardown. Off: these return 403, and the Studio hides endpoint management. |
--allowed-vertex-endpoints <id,...> (DGEM_ALLOWED_VERTEX_ENDPOINTS) |
none | Extra Vertex endpoints a request may select with vertex_url / X-DGem-Vertex-Url. The configured endpoint is always allowed; any other returns 400. |
Requests naming an unknown backend, or one that is not enabled, get 400 with the list of available backends
(HTTP header, query, JSON body and MCP alike); a request without a backend uses the gateway default. The Studio's
backend menu is per browser and sends X-DGem-Backend with each request; only an admin-enabled gateway lets it
change the default for everyone.
Keep the admin API off on shared gateways: with it on, any user who passes IAP can change routing for all users
or deploy and tear down the Vertex endpoint with the gateway's service account. Manage endpoints with
deploy_vertex_endpoint.sh instead, or run a separate admin-only gateway.
| Surface | How |
|---|---|
| Decision Studio | Topbar Backend Target menu (also shows live Vertex replica state) |
| HTTP API | Header X-DGem-Backend: vertex_first|vertex|cloudrun|local, query ?backend=..., or JSON "backend": "..." |
| MCP | "backend" argument on decide_policy, decide_custom_questions, locate_bounding_boxes |
| CLI through a gateway | -u https://<your-dgem-gateway>/v1 --gcp-auth (gateway default applies) |
| Gateway default | --default-backend / DGEM_DEFAULT_BACKEND |
Every response reports where it ran: header X-DGem-Backend-Used: vertex|cloudrun|local (and "backend_target" in
JSON bodies, including MCP tool results).
curl -sS https://<your-dgem-gateway>/api/decide/support_triage \
-H "Authorization: Bearer $(gcloud auth print-identity-token)" \
-H "Content-Type: application/json" -H "X-DGem-Backend: vertex_first" \
-d '{"variables": {"ticket": "Production database unreachable after certificate rotation."}}' -i \
| grep -i x-dgem-backend-used./bin/dgem serve --gcp-auth --vertex-url <ENDPOINT_ID> -u https://<CLOUD_RUN_URL>/v1 # both backends
./bin/dgem serve --local # laptop engine onlyGCP_PROJECT=<PROJECT> GCP_REGION=<REGION> GCP_PROJECT_NUMBER=<PROJECT_NUMBER> \
ALLOW_GROUP=<GROUP> \
DGEM_VERTEX_URL=<ENDPOINT_ID> DGEM_VERTEX_MODEL_ID=<MODEL_ID> \
DGEM_VERTEX_SA=dgemma-gpu-sa@<PROJECT>.iam.gserviceaccount.com \
DGEM_GATEWAY_HOSTS=<CUSTOM_DOMAIN> \
GATEWAY_TAG=candidate \
./scripts/deploy_cloudrun_gateway.sh- Builds the gateway image (Go binary + Studio) with Cloud Build, tagged with the git commit, and deploys
dgemma-gatewaywith--min-instances=1 --no-cpu-throttling, so its 4-minute Vertex health keepalive keeps running. ALLOW_GROUP(required) gets access through IAP / Cloud Run IAM;DGEM_VERTEX_URLis required unlessALLOW_NO_VERTEX=1(Cloud Run only).DGEM_GATEWAY_HOSTS: extra hostnames (for example a custom domain) the gateway treats as its own for IAM token handling.*.run.appand Vertex endpoints are recognised automatically.DGEM_BACKENDS,DGEM_ALLOWED_VERTEX_ENDPOINTSandDGEM_ADMIN_API=1are passed through when set (restrictions).GATEWAY_TAG: deploy as a tagged revision with no traffic, verify it athttps://<tag>---<gateway-host>, thengcloud run services update-traffic dgemma-gateway --to-tags=<tag>=100. Roll back by moving traffic to the previous revision.
Verify a new gateway revision before moving traffic:
GET /healthandGET /api/backend-config(expectedavailable_backends, Vertexstate: deployed).- One
/api/decidecall per backend, checkingX-DGem-Backend-Used. - MCP
tools/listand theget_health_and_gpu_statustool. - After moving traffic, repeat through the custom domain (without a token IAP should redirect to sign-in).
- Browser: sign in through IAP as a member of
<GROUP>. - CLI / scripts:
--gcp-authsends a Google ID token from Application Default Credentials. - MCP clients:
dgem mcp --remote https://<your-dgem-gateway>/mcpbridges stdio to the gateway with automatic ADC auth. Other options: Studio, MCP and HTTP API. - Observability: every gateway decision is traced (
dgem.gateway.decideand child spans); see Observability.
Latency and capacity and the operations runbook.