Skip to content

Latest commit

Β 

History

252 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

dgem β€” DiffusionGemma as a Zero-Shot Decision Model

dgem turns Google DeepMind's DiffusionGemma (26B-A4B-it, a discrete-diffusion Gemma 4) into a decision engine. You describe a decision as a small template of typed questions (boolean, choice, score). dgem places every question on the model's bidirectional canvas and reads a probability for every allowed answer, for every question, in one forward pass. There is no free-text generation to parse, and every answer comes with a per-question uncertainty score you can use to decide when to act automatically and when to escalate.

πŸ“– Docs site: ghchinoy.github.io/dgem

What's in this repo

Piece What it does Start here
Decision engine (dgem decide) Compiles .json.tmpl Policy-as-Template files into one-pass multi-question readouts with per-option probabilities and Shannon entropy. The Journey to Decision Models Β· Template Catalog
Confidence and calibration Per-answer probabilities and hesitation from one pass, calibration checks on labelled data, and research on option-order bias (Invariant Decision Calibration, IDC). Confidence and calibration Β· IDC
Entropy-gated cascade Answers low-uncertainty decisions directly and escalates the rest to Vertex AI gemini-3.8-flash with the Stage-1 probabilities attached. EXP-05
Four surfaces CLI, HTTP gateway (/api/decide, /v1/systemone), MCP server (dgem mcp, /mcp), and the embedded Decision Studio web app (dgem serve). Studio, MCP & API Β· CLI reference
Benchmarks & research log 9 reproducible dgem bench-* harnesses with committed JSON receipts, an experiment ledger (EXP-01–EXP-18), and a register of pre-registered follow-up experiments. Experiment Ledger Β· Proposed Experiments
Serving Vertex AI Dedicated Endpoint on RTX PRO 6000 (primary), Cloud Run GPU (scale-to-zero failover), GCE VMs, and local Apple Silicon (Metal). From laptop to production

Quick Start

Full walkthrough: Run on your laptop.

1. Start an Inference Backend

Choose the option matching your hardware:

  • Local Apple Silicon Mac (Metal) β€” No Docker or cloud needed:
    make setup && make download && make local-up   # engine on :8080, Decision Studio on :8090
  • Any Linux/Windows Workstation with NVIDIA GPU (Docker):
    # Lean public image (downloads the public weights on first boot); tags and digests: docs/deploy/public-images.md
    docker run --gpus all -p 8080:8080 us-central1-docker.pkg.dev/dgem-diffusiongemma/dgem/dgem:v0.1.0@sha256:5fa4a866163169aaf91c3bdb727020ff3c86e26e84ffadad86bc859053d45b26
  • Remote Google Cloud Endpoint:
    export DGEM_VERTEX_URL="<endpoint-id>" GCP_PROJECT_NUMBER="<project-number>" DGEM_GCP_AUTH=1

2. Build CLI & Execute Your First Decision

make build                                   # compiles ./bin/dgem
./bin/dgem decide -t templates/support_triage.json.tmpl \
  -v 'ticket=I was billed $500 twice for my annual renewal this morning!' --stats

3. Open Decision Studio (Web UI)

./bin/dgem serve --port 8090

Open http://localhost:8090 to inspect all 26+ templates, evaluate Stage 2 Gemini cascades, and visualize OpenTelemetry trace waterfalls.

Next, pick your path: build, deploy and operate Β· confidence and calibration Β· write your first policy.


Confidence Beyond Shannon: Invariant Decision Calibration (IDC)

In one sentence: IDC makes a dgem confidence score reflect the question, not the position where each answer was listed, and flags decisions whose answer depends on the list order. Plain-English guide with a worked example: docs/confidence-beyond-shannon.md.

The problem. Per-question Shannon entropy ($H = -\sum p_k \ln p_k$) is a useful escalation signal: it rises when human annotators disagree. But DiffusionGemma, like other language models, has a strong ballot-order ("Box A") habit: on blank, content-free questions it puts $88.3%$ / $78.3%$ / $49.3%$ of its probability on the first slot (2 / 3 / 4 options). On a borderline input that habit can make a toss-up look 99.9% certain, and the entropy gate waves it through.

What IDC does.

  1. Null-Prior De-Biasing (--null-prior-debias, no labeled data): divides out the measured slot habit. On the 50-item calibration suite it improved Brier from 0.175–0.193 (three same-session baselines) to 0.147; on the 231-item JevBench set it did not help (186 vs 187 correct, worse calibration). Suite-dependent, so validate before enabling.
  2. Dual-Mirror Canvas (--dual-mirror): adds a reversed-order copy of each choice question to the same canvas, so the forward and reversed readings come from one forward pass, and their gap (Mirror TVD) flags order-dependent answers. A research diagnostic: with lettered options the reversed slot lowers forward accuracy (letter collision, EXP-15); digit-labelled mirrors (--mirror-mode reversed-digits) stay within noise but add little error detection beyond hesitation (EXP-15–EXP-17). Not recommended in production.
  3. Slot Temperature Scaling (EXP-11): softens over-sharp scores. Needs labeled data. Fitted on held-out folds it cut ECE by 24–33% on 231 JevBench items ($T^* \approx 1.5$) but gave no reliable gain on the 50-item suite.

What's new. Removing a content-free prior (contextual calibration, Zhao et al. 2021), permutation debiasing (e.g. PriDe, Zheng et al. 2023), and temperature scaling (Guo et al. 2017) are known techniques. The part specific to a diffusion decision model is checking a reversed ballot on every request without a second forward pass, which turns order sensitivity from an offline audit into a per-request signal.

Status. A same-session re-run on 50 + 231 items (EXP-14, versioned receipts in benchmarks/runs/) and follow-ups (EXP-15–EXP-17) gave mixed results: the order-bias problem is real and reproducible, but the corrections are suite-dependent, and hesitation gating remains the recommended production signal. IDC is CLI-only today. See IDC Β§6 for every number and Proposed Experiments for what comes next.


Headline Results (measured, with sample sizes)

Result Value Sample Receipt
Single-pass accuracy, 11 public datasets (dgem bench-calibration) 88.0% (44/50) 50 benchmarks/results_calibration_cloudrun.json
+ Null-prior de-biasing (IDC), same session 45/50, Brier 0.147 vs 0.175–0.193 (3 baselines) 50 benchmarks/runs/20260925-vertex-idc/
+ Entropy cascade to gemini-3.8-flash ($\tilde H \ge 0.16$, 34% escalated) 98.0% (49/50), 56% lower cost than Gemini on every item 50 results_calibration_cascade_normalized.json
JevBench v1.3.1: dgem single pass β†’ entropy cascade (hesitation β‰₯ 16%, 39% escalated, offline) 187–189 β†’ 221 of 231 231 benchmarks/runs/20260925-vertex-idc/
JevBench v1.3.1: null-prior de-biasing 186 of 231, Brier 0.293 vs 0.264 (no gain) 231 benchmarks/runs/20260925-vertex-idc/
Decision Index panel: bracket routing (> 26 options) + slot batching 76.67 β†’ 98.89, coverage 16/22 β†’ 22/22 22 requests benchmarks/decision_index/
Listwise reranking of 10 passages in one pass (EXP-10) 0.9265 nDCG@10, 0% ties 30 queries results_rerank_cloudrun.json
Content-free Slot-A habit (EXP-13B) 88.3% / 78.3% / 49.3% for K = 2 / 3 / 4 probe results_permutation_cloudrun.json

Run-to-run noise is about Β±1 item on 50 and Β±2 on 231. Cascade thresholds were chosen on the evaluation items; treat these as directional. Compare runs with python3 scripts/bench_runs.py compare. Details and caveats: Benchmark Report, Experiment Ledger.


Four Ways to Use dgem

See Decision Studio, MCP and HTTP API and Gateway and routing for full details:

Interaction Surface Command / Endpoint Description
1. πŸ–₯️ Decision Studio Web App ./bin/dgem serve --port 8090 Embedded Lit WebComponents web application featuring all 26+ .json.tmpl decision policies (core, calibration, multimodal, rerank), topbar Backend Target selector (vertex_first | vertex | cloudrun), Stage 2 Gemini Cascade (gemini-3.8-flash), live SigLIP 2D Bounding Box SVG overlays (EXP-09), a plain-English Concepts tab (including an IDC walkthrough), and OpenTelemetry Trace Waterfall inspection.
2. πŸ€– Model Context Protocol (MCP) ./bin/dgem mcp (stdio)
POST /mcp (Streamable HTTP)
Native MCP server exposing 6 tools (decide_policy, locate_bounding_boxes, decide_custom_questions, list_policy_templates, get_health_and_gpu_status, warmup_gpu) with backend (vertex_first | vertex | cloudrun) and Stage 2 Gemini Cascade support (cascade_mode, cascade_threshold, cascade_model).
3. 🌐 HTTP Gateway REST API POST /api/decide/{template}
POST /v1/systemone, GET /api/templates
Execute any .json.tmpl decision policy or /v1/systemone schema with X-DGem-Backend: vertex_first | vertex | cloudrun (X-DGem-Backend-Used returned on every response) and optional Stage 2 gemini-3.8-flash cascade.
4. ⌨️ CLI & 9 Benchmark Harnesses ./bin/dgem decide --vertex-url ...
./bin/dgem bench-*
Direct single-pass decisions (--stats, --null-prior-debias, --dual-mirror) and nine reproducible evaluation harnesses (bench, bench-ecotone, bench-intents, bench-calibration, bench-bbox, bench-rerank, bench-jev, bench-decision-index, bench-permutation) backed by docs/experiments/ (EXP-01 – EXP-13).

Why a "Decision Model"?

Production teams have historically chosen between two extremes for automated triage, routing, and guardrails:

  1. Discriminative classifiers & automata (BERT / DeBERTa / C++ WFSTs): very fast, but rigid. A new policy rule or category means new labeled data, retraining, and redeployment.
  2. Autoregressive LLMs (Gemini / GPT / Gemma 4): zero-shot flexible, but they generate answers token by token (seconds per multi-field JSON answer), can drift from the output format, and their token probabilities are spread across formatting tokens rather than the decision itself.

A zero-shot decision model sits in between. dgem compiles a .json.tmpl template into a fixed diffusion canvas (32–256 tokens) with bidirectional attention; boolean gates, [A–Z] choices, and ordinal scores are read together in one forward pass, and every answer is constrained to the allowed options.

Architectural Dimension Discrete Diffusion Decision Model (dgem) Discriminative Encoder (DeBERTa-v3 / Llama-Guard) Autoregressive LLM (Gemini / Gemma 4) Compiled Rulebook (ecotone C++ WFST)
Policy Adaptability Zero-shot Policy-as-Template (edit .json.tmpl) Labeled dataset & retraining per label change Zero-shot prompt engineering Manual grammar authoring & compilation
Inference Latency ~55 ms GPU / ~150 ms end to end for a 3-question decision on RTX PRO 6000; ~1.4 s for a 12-slot rerank (measured on L4) ~5 – 25 ms (single head) 17,486.6 ms (~17.5 s for 3-slot JSON + CoT) 1.35 – 8.68 ms (1.54 ms p50 over UDS)
Passes per Request 1 forward pass for all questions (cost grows with canvas length) One classifier per attribute One token per step ($O(T_{\text{output}})$) $O(N_{\text{chars}})$ graph traversal
Joint Slot Conditioning Bidirectional (slot_1 <-> slot_2) in a single pass Independent heads Left-to-right only Local sliding window (1–3 tokens)
Uncertainty & Calibration Per-option probabilities + entropy; label-free order-bias correction (null-prior) and same-pass reversed-ballot check (IDC) Often overconfident out-of-distribution Sequence-level logprobs over formatting tokens Static arc weights
Guardrail Examples (50-item suite) AgentDrift 7/7, prompt injection 4/4, RAG grounding 2/2 Narrow single-task scope High accuracy, 15–25Γ— higher latency 36.7% on semiotic polysemy traps

Serving: From Laptop to Production

Measured with the current serving image (3-question decision, samples: 1, p50; receipts):

Stage Where Cold start GPU / end to end Guide
Crawl Apple Silicon (diffgemma) or a local NVIDIA GPU None ~0.9 s on Metal Run on your laptop
Walk Cloud Run GPU, 1Γ— RTX PRO 6000, scale to zero 2.3–2.7 min 61 / 153 ms Deploy on Cloud Run
Run Vertex AI dedicated endpoint, g4-standard-48 + 1Γ— RTX PRO 6000 None 55 / 150 ms Production on Vertex AI

A gateway (dgem serve) in front routes to Vertex first and fails over to Cloud Run (Gateway and routing). Capacity, runbook and observability: Latency and capacity Β· Operations runbook.


Installation & CLI Examples

# Clone the repository
git clone https://github.com/ghchinoy/dgem.git
cd dgem

# Compile dgem binary into bin/
make build

1. Single-Pass Discrete Decision (dgem decide)

Evaluate customer tickets, code changes, or security alerts in a single sub-second forward pass:

./bin/dgem decide -t templates/support_triage.json.tmpl \
  -v 'ticket=I was billed $500 twice for my annual renewal this morning!' \
  --stats

Output:

QUESTION         | TYPE       | VALUE / CHOICE       | CONFIDENCE | ENTROPY (H) | AGREEMENT 
-----------------------------------------------------------------------------------------
sentiment        | score      | frustrated           | 99.8%      | 0.002 nats  | 1.00      
team             | choice     | billing              | 100.0%     | 0.000 nats  | 1.00      
urgent           | boolean    | yes                  | 99.9%      | 0.001 nats  | 1.00      

──────────────────────────────── STATS ────────────────────────────────
  Model:             nvidia/diffusiongemma-26B-A4B-it-NVFP4
  Endpoint:          http://127.0.0.1:8080/v1/chat/completions
  Total Wall Time:   856 ms
  KV Cache Reused:   169 tokens (82.8% hit rate)
  Denoise Steps:     1 step (policy: samples=4)
───────────────────────────────────────────────────────────────────────

2. Generative Prompt Completion (dgem ask)

Standard chat completion with optional thinking mode:

./bin/dgem ask "Explain discrete block diffusion in two sentences."

3. Remote Cloud Routing with IAM Authentication

Connect to any remote GCE or Cloud Run GPU service:

./bin/dgem decide \
  -u "http://<EXTERNAL_IP>:8080/v1" \
  -m "nvidia/diffusiongemma-26B-A4B-it-NVFP4" \
  -t templates/support_triage.json.tmpl \
  -v 'ticket=Outage: production database cluster unreachable' \
  --stats

4. Multimodal Visual Assessment (--image / -I)

Attach local image paths (automatically base64 encoded) or remote URLs:

./bin/dgem decide -t templates/multimodal/ui_design_review.json.tmpl \
  -I fixtures/ui_component.svg \
  -v 'component=CheckoutCard' \
  --stats

Benchmark Suites & Empirical Calibration

dgem includes nine benchmark harnesses (all tracked in docs/experiments/README.md). Four of the most commonly used are below; the others are bench-jev (JevBench v1.3.1), bench-decision-index (Decision Index panel + /v1/systemone), bench-permutation (option-order sensitivity and IDC, EXP-13), bench-rerank (listwise reranking, EXP-10), and bench-bbox (bounding boxes, EXP-09).

1. Public Dataset Policy & Epistemic Calibration Suite (dgem bench-calibration)

Evaluates 50 items across 11 public datasets (benchmarks/calibration_suite.jsonl), testing declarative policy templates (templates/calibration/*.json.tmpl) across agent trajectory hijacking (AgentDrift), multilingual jailbreaks (deepset/prompt-injections), RAG fact grounding (LLM-AggreFact), retrieval relevance (MS MARCO), toxicity (Jigsaw Civil Comments), and human annotator disagreement (ChaosNLI):

./bin/dgem bench-calibration -u "https://<CLOUD_RUN_URL>/v1" --gcp-auth -w 4 \
  -o benchmarks/results_calibration_cloudrun.json
Public Dataset / Policy Domain Cases Accuracy Mean $P(y)$ Mean Entropy $H$ Avg Latency
AgentDrift (agent_step_drift.json.tmpl β€” Hijack + 4-Way Step Localization) 7 100.0% (7/7) ⭐ 0.997 0.0186 nats 693 ms
deepset/prompt-injections (prompt_injection.json.tmpl β€” en/de Gate) 4 100.0% (4/4) ⭐ 0.980 0.0817 nats 669 ms
LLM-AggreFact & MS MARCO (RAG Grounding & Retrieval Relevance) 4 100.0% (4/4) ⭐ 0.993 0.0403 nats 728 ms
CLINC150, Banking77, GoEmotions, BoolQ, Yelp/SST-5 20 100.0% (20/20) ⭐ 0.898 0.3263 nats 769 ms
ChaosNLI Crowd Consensus (low-entropy) 3 100.0% (3/3) 0.986 0.0744 nats (1.0Γ—) 625 ms
ChaosNLI Crowd Split (high-entropy) 3 33.3% (1/3) 0.759 0.5932 nats (8.0Γ— higher; n=3) 731 ms
Stage 1 Alone: DiffusionGemma (steps=1, think=0) 50 88.0% (44/50) 0.925 0.2279 nats 712 ms
Raw Entropy Cascade (EXP-05a): dgemma [H<0.35] $\rightarrow$ gemini-3.8-flash 50 94.0% (47/50, +6.0%) 0.959 0.1410 nats 1,824 ms (72% early-exit)
Normalized + Prior-Guided Cascade (EXP-05b, $\tilde{H} &lt; 0.16$, threshold tuned on these items) 50 98.0% (49/50, +10.0%) ⭐ 0.960 0.1416 nats ($\tilde{H}=0.106$) 2,105 ms (66% early-exit)
Stage 2 Alone: gemini-3.8-flash (100% Frontier LLM) 50 98.0% (49/50) 0.959 0.1347 nats 3,412 ms (4.8Γ— slower)

2. Multi-Domain Operational Triage (dgem bench)

Evaluates 30 multi-field test cases (boolean + choice + score in a single pass) across support, code_review, and security (benchmarks/eval_dataset.jsonl):

./bin/dgem bench -d benchmarks/eval_dataset.jsonl -M slot -o benchmarks/results_cloudrun.json

3. Ecotone WFST vs. DiffusionGemma (dgem bench-ecotone)

Evaluates 49 Text Normalization cases comparing C++ ecotone (OpenFst / Sparrowhawk WFSTs over unix:///tmp/ecotone.sock) against DiffusionGemma across semiotic polysemy traps and deterministic NSWs:

./bin/dgem bench-ecotone -c benchmarks/ecotone/tn_semiotics.jsonl --samples 1 -o benchmarks/results_ecotone.json

4. High-Cardinality Intent & Out-of-Scope Routing (dgem bench-intents)

Evaluates 30-way to 151-way intent routing and Out-of-Scope (oos) rejection on PolyAI/banking77 and DeepPavlov/clinc150:

./bin/dgem bench-intents --dataset banking77 --full --workers 16
./bin/dgem bench-intents --dataset clinc150 --full --workers 16

Documentation

Published at ghchinoy.github.io/dgem (index), organized by audience:


Contributing

Issues, bug reports, and feature discussions are welcome! However, we are not accepting pull requests (PRs) at this time. If you encounter a bug or have feedback on benchmark methodologies or templates, please open an Issue.

License

This project is licensed under the Apache-2.0 License.

Disclaimer

Caution

This is not an officially supported Google product. This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.

About

Decision Model - Calibrated DiffusionGemma template evaluation cli, mcp, web app for fast, confident zero-shot decision making.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages