Skip to content

Latest commit

 

History

History
296 lines (242 loc) · 22.2 KB

File metadata and controls

296 lines (242 loc) · 22.2 KB
title Decision Studio Web App, MCP Server & HTTP Gateway API
description Complete guide to the four interaction surfaces of dgem: the interactive Lit WebComponents Decision Studio (dgem serve), the Model Context Protocol (MCP) Server for AI Agents (dgem mcp & POST /mcp), the HTTP Gateway REST API, and the CLI.

Decision Studio Web App, MCP Server & HTTP Gateway API

Beyond the command-line evaluation harness (dgem decide, dgem bench-*), dgem ships as a unified HTTP API Gateway, Lit WebComponents Decision Studio, Scale-to-Zero GPU Cold-Start Orchestrator, and Model Context Protocol (MCP) Server in a single self-contained Go binary (studio.DistFS() embedded).


1. The Four Ways to Use dgem

flowchart LR
    subgraph Clients["Four Ways to Interact with dgem"]
        UI["1. Decision Studio Web App\nBrowser UI (Lit WebComponents)\nhttp://localhost:8090/"]
        MCP["2. AI Agents via MCP\nGemini CLI · Claude · Cursor\ndgem mcp (stdio) & POST /mcp"]
        REST["3. Microservices & curl\nSimplified REST JSON API\nPOST /api/decide/{template}"]
        CLI["4. Terminal & CI/CD\nGo CLI & Benchmark Suites\ndgem decide / dgem bench-*"]
    end

    subgraph Gateway["dgem serve / dgem mcp Gateway"]
        CAT["26+ .json.tmpl Policy Catalog\n(core · calibration · multimodal · rerank)"]
        WARM["Scale-to-Zero GPU Cold-Start\nOrchestrator & IAM/IAP Token Minter"]
        OTEL["OpenTelemetry & Cloud Trace\nSpan Waterfall Ring Buffer"]
    end

    subgraph Backends["DiffusionGemma GPU Backends"]
        CR["Serverless Cloud Run GPU\n1× RTX Pro 6000 (48GB) / 1× L4 (24GB)"]
        METAL["Local Apple Silicon Metal\ndiffgemma serve (:8080)"]
    end

    UI --> Gateway
    MCP --> Gateway
    REST --> Gateway
    CLI --> Gateway
    Gateway --> CR
    Gateway --> METAL
Loading
Interaction Surface Command / Endpoint Target Audience Key Capabilities
1. Decision Studio Web App dgem serve $\rightarrow$ http://localhost:8090/ Engineers, PMs, Security & AI Researchers Interactive Lit WebComponents playground with live .json.tmpl policy catalog (core, calibration, multimodal, rerank), SigLIP bounding-box SVG canvas (EXP-09), one-click scale-to-zero GPU warmup, and OpenTelemetry span waterfall viewer.
2. Model Context Protocol (MCP) dgem mcp (stdio) or POST /mcp (Streamable HTTP) AI Agents (Gemini CLI, Claude Desktop, Cursor, Antigravity) Exposes 6 native MCP tools (decide_policy, locate_bounding_boxes, decide_custom_questions, list_policy_templates, get_health_and_gpu_status, warmup_gpu) over both stdio and stateless Streamable HTTP.
3. HTTP Gateway REST API POST /api/decide/{template} & POST /v1/chat/completions Microservices, Web Backends, curl / Python scripts Execute named .json.tmpl policies with a simple JSON variable map—no local dgem CLI or .json.tmpl files required by the caller. Automatically mints GCP IAM/IAP tokens and holds requests while scale-to-zero Cloud Run GPUs wake up.
4. dgem CLI & Benchmarks dgem decide, dgem ask, dgem bench-* Terminal workflows, CI/CD pipelines, Reproducible research Direct single-pass decisions (--stats) and full evaluation harnesses (bench, bench-ecotone, bench-intents, bench-calibration, bench-bbox, bench-rerank).

2. Decision Studio Web App (dgem serve)

Launching dgem serve starts the embedded Lit WebComponents Decision Studio alongside the REST API Gateway and Streamable HTTP MCP Server:

# Launch Decision Studio locally on port 8090 pointing at a Serverless Cloud Run GPU backend
./bin/dgem serve \
  -u "https://<CLOUD_RUN_URL>/v1" \
  --gcp-auth \
  --port 8090

Open http://localhost:8090 in your browser. Decision Studio provides four integrated workspaces:

2.1 Live Policy Catalog & One-Click Preset Execution

  • Automatically discovers all 26+ .json.tmpl decision policies across four categories:
    • core: support_triage, code_review, security_incident, contract_audit, agent_router, secops_conditional_dag (EXP-07), tn_disambiguation (EXP-02), intent_banking77 / intent_clinc150 (EXP-03).
    • calibration: agent_step_drift, prompt_injection, grounding_claim_check, chaos_nli, ms_marco_relevance, civil_comments_toxicity, safety_toxicity, sentiment_emotion (EXP-04 & EXP-05).
    • multimodal: bbox_localization, bbox_multi_object_detr, bbox_multi_object_set (EXP-09 SigLIP spatial grounding).
    • rerank: listwise_decision_rerank (12 simultaneous slots: 10 passages + 2 RAG security/abstention gates) and pointwise_rerank (EXP-10).
  • Selecting any policy automatically populates curated sample variables (ticket, diff, event_log, query, candidates, claim, context) and displays the live Go .json.tmpl source so users can test or customize policies in real time.

2.2 Answer Cards: Probabilities & Hesitation

For every question evaluated in the single forward pass (steps=1, think=0), Decision Studio shows:

  • The chosen value and its probability $p_i = \exp(\text{logprob})$, plus bars for the other allowed answers.
  • A Hesitation score from 0% (all probability on one answer) to 100% (a perfect tie): normalized Shannon entropy $\tilde{H} = H / \ln K$, so 2-option and 26-option questions share one scale. It is labelled Clear (under 16%, matching the EXP-05b cascade gate), Somewhat unsure (16–50%), or Very unsure (above 50%). Hover to see the raw entropy in nats.
  • A summary strip: Model passes, Refinement steps, Response time, and Highest hesitation across the request.

Low hesitation is a strong signal but not a guarantee; option order can hide doubt. See Confidence Beyond Shannon (IDC). The helper that computes hesitation lives in studio/src/hesitation.ts.

2.2b Concepts Walkthrough (Plain-Language)

The Concepts tab is a self-serve, jargon-light introduction with five tabs: What's a decision model?, One pass vs. word-by-word, When to trust an answer (hesitation + an interactive IDC order-check demo), Built-in guardrails, and a Glossary (basics first, technical names collapsed). Simulations are labelled as illustrative. A Presenter mode toggle shows an optional talk track for live demos. A standalone copy (without the IDC demo and glossary) is published at visualizer.html.

2.3 Interactive Multimodal SigLIP Bounding Box Canvas (EXP-09)

When a multimodal/* policy (bbox_localization, bbox_multi_object_detr) is selected:

  • Users can pick from built-in fixtures/bbox/*.png test scenes or drag-and-drop custom images.
  • Decision Studio draws both the Continuous Softmax Expectation Box ($\hat{c}m = \sum_k v_k p{m,k}$, green solid overlay) and the Discrete 21-Bin Argmax Box (amber dashed overlay) directly over the image, alongside Per-Edge Occlusion Entropy bars (ymin, xmin, ymax, xmax).

2.4 Scale-to-Zero GPU Cold-Start Orchestrator & OpenTelemetry Waterfall

  • Live GPU State Pill: Polls GET /api/status to show whether the upstream Cloud Run GPU (NVIDIA RTX Pro 6000 48GB or NVIDIA L4 24GB) is warm_and_ready, warming_up (with elapsed seconds and progress bar), or scaled_to_zero.
  • One-Click GPU Wakeup (POST /api/warmup): Wakes a 0-instance Cloud Run GPU in the background and streams readiness to all connected browser tabs and MCP agents simultaneously (MarkGPUWarm() broadcast channel).
  • OpenTelemetry Trace Inspector (GET /api/traces): Displays the exact W3C trace_id, splitting total wall time into cold_start_wait_ms (GCS FUSE weight streaming) vs. gpu_forward_ms (pure DiffusionGemma forward pass), and exports spans to Google Cloud Trace when GOOGLE_CLOUD_PROJECT is set.

3. Model Context Protocol (MCP) Server (dgem mcp & POST /mcp)

dgem includes a native Model Context Protocol (MCP) server built with the official Go MCP SDK (github.com/modelcontextprotocol/go-sdk/mcp). It supports two transports simultaneously:

  1. Standard Input/Output (stdio) via dgem mcp: Ideal for local AI coding assistants (Gemini CLI, Claude Desktop, Cursor, Antigravity).
  2. Streamable HTTP via POST /mcp on dgem serve: Ideal for remote Cloud Run deployments and multi-agent cloud orchestrators (Stateless: true, JSONResponse: true).

3.1 Configuring dgem mcp (Automatic ADC Auth) in Antigravity (~/.gemini/config/mcp_config.json), Gemini CLI (~/.gemini/settings.json), or Claude Desktop

dgem mcp automatically reads Application Default Credentials (ADC: ~/.config/gcloud/application_default_credentials.json) in pure Go to mint and cache both OIDC id_tokens (accepted by Cloud Run IAP programmaticClients on your gateway) and OAuth2 access_tokens (for Vertex AI Dedicated Endpoint <endpoint-id>).

Option A: Connect to a Remote dgem serve Gateway via ADC (--remote)

{
  "mcpServers": {
    "dgem-remote": {
      "command": "/path/to/dgem/bin/dgem",
      "args": [
        "mcp",
        "--remote",
        "https://<your-dgem-gateway>/mcp"
      ]
    }
  }
}

Option B: Local stdio with Automatic ADC & vertex_first Routing

{
  "mcpServers": {
    "dgem": {
      "command": "/path/to/dgem/bin/dgem",
      "args": [
        "mcp"
      ],
      "env": {
        "DGEM_VERTEX_URL": "<ENDPOINT_ID>",
        "GCP_PROJECT_NUMBER": "<PROJECT_NUMBER>",
        "DGEM_REMOTE_URL": "https://<CLOUD_RUN_URL>/v1"
      }
    }
  }
}

The MCP server runs inside your agent and calls the endpoints directly with ADC tokens. It needs to know where they are: DGEM_VERTEX_URL (with GCP_PROJECT_NUMBER for a bare endpoint ID) for Vertex AI, and DGEM_REMOTE_URL (or -u) for the Cloud Run /v1 URL used by cloudrun and vertex_first failover. Without them, requests go to http://127.0.0.1:8080/v1. If a gateway is already deployed, Option A is simpler.

Option C: Direct Streamable HTTP (dgem serve on localhost)

{
  "mcpServers": {
    "dgem-local-http": {
      "serverUrl": "http://localhost:8090/mcp"
    }
  }
}

Option D: Local stdio Targeting Apple Silicon Metal (--local)

{
  "mcpServers": {
    "dgem-local": {
      "command": "/path/to/dgem/bin/dgem",
      "args": [
        "mcp",
        "--local"
      ]
    }
  }
}

--local sends every call to the diffgemma engine at http://127.0.0.1:8080/v1 (change it with -u), with no cloud credentials or fallback. Start the engine first (make serve); see Run on a laptop.

opencode

opencode uses its own config format (~/.config/opencode/opencode.json or a project opencode.json). Any of the options above works; for example, Option D:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "dgem-local": {
      "type": "local",
      "command": ["/path/to/dgem/bin/dgem", "mcp", "--local"],
      "enabled": true
    }
  }
}

For Option A use "command": ["/path/to/dgem/bin/dgem", "mcp", "--remote", "https://<your-dgem-gateway>/mcp"]. Tools appear in opencode as dgem-local_decide_custom_questions and so on. Restart opencode after rebuilding bin/dgem: each session keeps its own dgem mcp process.

3.2 Complete Catalog of MCP Tools Exposed by dgem

MCP Tool Name Input Arguments Output Payload & Purpose
get_health_and_gpu_status {"backend": "vertex_first|vertex|cloudrun|local"} (optional) Returns gateway_healthy, gpu_available, gpu_state (warm_and_ready, warming_up, scaled_to_zero), active_backend, seconds_since_last_read, estimated_wake_seconds, gpu_tier, and templates_available for the selected backend (the server default if omitted). Agents call this first to check if the GPU is warm.
warmup_gpu {"wait_for_ready": true | false} Triggers a scale-from-zero wakeup (0 -> 1 instance) on the Cloud Run GPU backend. Set wait_for_ready: false to start streaming the 17.53 GiB safetensors over GCS FUSE asynchronously while the agent performs other work.
list_policy_templates {"category": "core" | "calibration" | "multimodal" | "rerank"} (optional) Lists the executable .json.tmpl decision policies found in the templates directory, with their required variables and sample_vars.
decide_policy {"template": "support_triage", "variables": {...}, "image": "...", "backend": "vertex_first|vertex|cloudrun|local", "vertex_url": "...", "cascade_mode": "off|entropy|on_miss", "cascade_threshold": 0.35, "cascade_model": "gemini-3.8-flash", "expected_answers": {...}, "suggest_expansions": false, "expansion_entropy": 0.35} Executes any named .json.tmpl policy in $O(1)$ forward passes on vertex_first (default), vertex, cloudrun, or local (Apple Silicon Metal), with optional Stage 2 Gemini Cascade (gemini-3.8-flash) and taxonomy expansion. Returns answers (restricted-softmax probabilities per slot), diagnostics (per-slot entropy), and the summary fields in 3.3.
decide_custom_questions {"context": "...", "questions": [{"id": "...", "type": "boolean|choice|score", "question": "...", "options": [...]}], "backend": "vertex_first|vertex|cloudrun|local", "vertex_url": "...", "cascade_mode": "off|entropy|on_miss", "cascade_threshold": 0.35, "cascade_model": "gemini-3.8-flash", "expected_answers": {...}, "suggest_expansions": false, "expansion_entropy": 0.35} Evaluates an ad-hoc multi-slot decision schema dynamically constructed by the calling agent in 1 forward pass—the MCP equivalent of POST /v1/systemone—without needing a .json.tmpl file on disk. options is required for choice slots (at most 26). Returns the same fields as decide_policy.
locate_bounding_boxes {"image": "<url-or-data-uri>", "target": "checkout button", "backend": "vertex_first|vertex|cloudrun", "vertex_url": "..."} Executes single-pass SigLIP 2D spatial localization (EXP-09), returning both softmax_expectation_box_1000 ($\hat{c}m = \sum_k v_k p{m,k}$) and discrete_argmax_box_1000 [ymin, xmin, ymax, xmax] in [0, 1000] coordinates plus per-edge occlusion entropy (coordinate_entropy_nats). (Note: local Metal is text-only and does not support SigLIP vision).

3.3 Reading decide tool results

decide_policy and decide_custom_questions return the same top-level fields as POST /api/decide, apart from trace_id and trace_spans (MCP calls aren't traced through the gateway root span; see Observability).

Field Meaning
answers.<slot> The chosen value (label, plus choice or score), its confidence, the probabilities of every allowed answer, and agreement / stderr across samples.
diagnostics.questions.<slot>.entropy Shannon entropy of that slot in nats. Convert it to Hesitation % by dividing by $\ln K$ ($K$ = number of allowed answers); see 2.2.
max_entropy The highest slot entropy in the request (nats). Use it as a single "is anything uncertain?" signal, for example to decide whether to ask a person or escalate with cascade_mode: "entropy".
wall_time_ms Total time spent in the tool call.
gpu_forward_ms Time of the successful model call (the round trip to the engine).
cold_start_wait_ms Time spent waiting for a scaled-to-zero backend to wake (retries and backoff), 0 when warm.
warmup_attempts Number of attempts, including retries while waking.
backend_target, upstream_url, model Where the decision ran: vertex, cloudrun or local, and the upstream URL and model.
cascade, suggested_expansions Present only when cascade_mode or suggest_expansions is used.

decision repeats answers and diagnostics in the raw engine format, for clients that want the unmodified response.

3.4 Troubleshooting MCP

Symptom Cause and fix
Gemini-backed clients (opencode, Gemini CLI) fail on every request with functionDeclaration parameters.questions schema specified other fields alongside any_of The dgem binary predates the Gemini-safe tool schemas (issue #12). Rebuild (go build -o bin/dgem .) and restart every client session: each session keeps its own dgem mcp process, so old ones keep running the old binary. For a hosted gateway, redeploy it.
connection refused on 127.0.0.1:8080 with dgem mcp --local The local diffgemma engine isn't running. Start it with make serve and check curl -s http://127.0.0.1:8080/v1/models.
The first local call takes 15–40 seconds Expected on Apple Silicon: the first request after the engine starts pays for prefill (12–33 s in our runs). Later calls reuse the cached prompt prefix.
max_entropy or gpu_forward_ms is always 0 The binary or hosted gateway predates issue #14. Rebuild or redeploy.
Calls fail with HTTP 429 or take minutes A Cloud Run or Vertex backend is waking from zero. Call get_health_and_gpu_status, and warmup_gpu with wait_for_ready: false before a batch.
Plain dgem mcp calls go to 127.0.0.1:8080 No endpoints were configured; see Option B, or use --remote (Option A).
validating /properties/backend: enum: vertex does not equal any of: [local] (MCP), or backend "vertex" is not enabled on this gateway (available: local) (HTTP API, and MCP locate_bounding_boxes on a local-only server) The backend argument names a backend that isn't configured. It's no longer rerouted silently. Each MCP tool's schema lists the configured backends as an enum, so the MCP server rejects any other value before the call runs; the list in brackets is what this server offers. Omit backend to use the default, pick one from the list, or configure it: DGEM_VERTEX_URL / --vertex-url for vertex and vertex_first, -u for cloudrun. On a hosted gateway the operator may also restrict backends with DGEM_BACKENDS.
vertex_url "..." is not allowed on this gateway vertex_url may only name the configured endpoint, or one listed in --allowed-vertex-endpoints (DGEM_ALLOWED_VERTEX_ENDPOINTS) on dgem serve. Drop the argument to use the configured endpoint.

To check what a dgem binary advertises without an agent, pipe initialize and tools/list into it (the sleep keeps stdin open long enough for the reply; the server exits when stdin closes):

{ printf '%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"check","version":"1"}}}' \
  '{"jsonrpc":"2.0","method":"notifications/initialized"}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}'; sleep 2; } \
  | ./bin/dgem mcp --local | jq -c 'select(.id == 2) | .result.tools[].name'

It should list the six tools: decide_custom_questions, decide_policy, get_health_and_gpu_status, list_policy_templates, locate_bounding_boxes and warmup_gpu.


4. HTTP Gateway REST API Reference (dgem serve)

Any service or script can query dgem serve (https://<your-dgem-gateway> or local http://localhost:8090) using standard HTTP/JSON. Pass X-DGem-Backend: vertex_first | vertex | cloudrun | local (or ?backend=local / JSON "backend": "local") to select the GPU target; every response returns X-DGem-Backend-Used: vertex | cloudrun | local.

Endpoint Method Description
/api/decide & /api/decide/{template} POST Renders {template}.json.tmpl (or inline custom_template) with {"variables": {...}, "backend": "vertex_first", "cascade_mode": "off|entropy|on_miss", "cascade_threshold": 0.35, "cascade_model": "gemini-3.8-flash"}, runs 1-pass DiffusionGemma readout, and returns answers, diagnostics, backend_target, max_entropy, gpu_forward_ms, cold_start_wait_ms, and trace_spans.
/v1/systemone POST Direct pass-through proxy to structured_server.py's /v1/systemone (SystemOne / JevBench multipart image + JSON state/questions schema evaluation) across Vertex AI (/invoke/v1/systemone) or Cloud Run GPU (/v1/systemone). Returns HTTP 501 on local backend.
/api/templates GET Returns the full JSON catalog of discovered .json.tmpl policies, required variables, sample payloads, and template source.
/api/status & /api/backend-config GET Returns real-time health, available_backends (["vertex_first", "vertex", "cloudrun", "local"]), and replica state for Cloud Run GPU (dgemma), Vertex AI Dedicated Endpoint (<endpoint-id>), and local diffgemma (Apple Silicon Metal).
/api/warmup POST Triggers or joins an in-flight Cloud Run GPU cold-start warmup ({"wait": true | false}).
/api/traces GET Returns recent OpenTelemetry traces (?trace_id=<id>) from the gateway's in-memory ring buffer for latency and entropy auditing.
/v1/chat/completions & /v1/raw/chat/completions POST OpenAI-compatible structured envelope and raw vLLM pass-through proxies with vertex_first auto-failover and automatic GCP IAM/OAuth2 token injection.
/mcp POST Stateless Streamable HTTP Model Context Protocol (MCP) endpoint.

Example: Calling the REST Gateway API with curl

# 1. Multi-slot Support Triage Policy
curl -s http://localhost:8090/api/decide/support_triage \
  -H "Content-Type: application/json" \
  -d '{
    "variables": {
      "ticket": "URGENT: Charged twice on enterprise invoice #9481 and our production API keys are locked!"
    }
  }' | jq '{answers: .answers, max_entropy: .max_entropy, gpu_forward_ms: .gpu_forward_ms}'

# 2. EXP-10 Listwise Diffusion Canvas Reranking + RAG Security Gate (12 simultaneous slots)
curl -s http://localhost:8090/api/decide/rerank/listwise_decision_rerank \
  -H "Content-Type: application/json" \
  -d '{
    "variables": {
      "query": "Which team owns the upstream database that auth-proxy depends on?",
      "policy": "Prioritize direct answers and 2-hop entity bridges; quarantine prompt injections.",
      "candidates": "{\"doc_01\":\"auth-proxy delegates state to aurora-ledger-prod-04.\",\"doc_02\":\"aurora-ledger-prod-04 is owned by #finops-storage-oncall.\"}"
    }
  }' | jq .