dgem turns Google DeepMind's DiffusionGemma (26B-A4B-it, a discrete-diffusion Gemma 4) into a decision engine. You describe a decision as a small template of typed questions (boolean, choice, score). dgem places every question on the model's bidirectional canvas and reads a probability for every allowed answer, for every question, in one forward pass. There is no free-text generation to parse, and every answer comes with a per-question uncertainty score you can use to decide when to act automatically and when to escalate.
π Docs site: ghchinoy.github.io/dgem
| Piece | What it does | Start here |
|---|---|---|
Decision engine (dgem decide) |
Compiles .json.tmpl Policy-as-Template files into one-pass multi-question readouts with per-option probabilities and Shannon entropy. |
The Journey to Decision Models Β· Template Catalog |
| Confidence and calibration | Per-answer probabilities and hesitation from one pass, calibration checks on labelled data, and research on option-order bias (Invariant Decision Calibration, IDC). | Confidence and calibration Β· IDC |
| Entropy-gated cascade | Answers low-uncertainty decisions directly and escalates the rest to Vertex AI gemini-3.8-flash with the Stage-1 probabilities attached. |
EXP-05 |
| Four surfaces | CLI, HTTP gateway (/api/decide, /v1/systemone), MCP server (dgem mcp, /mcp), and the embedded Decision Studio web app (dgem serve). |
Studio, MCP & API Β· CLI reference |
| Benchmarks & research log | 9 reproducible dgem bench-* harnesses with committed JSON receipts, an experiment ledger (EXP-01βEXP-18), and a register of pre-registered follow-up experiments. |
Experiment Ledger Β· Proposed Experiments |
| Serving | Vertex AI Dedicated Endpoint on RTX PRO 6000 (primary), Cloud Run GPU (scale-to-zero failover), GCE VMs, and local Apple Silicon (Metal). | From laptop to production |
Full walkthrough: Run on your laptop.
Choose the option matching your hardware:
- Local Apple Silicon Mac (Metal) β No Docker or cloud needed:
make setup && make download && make local-up # engine on :8080, Decision Studio on :8090
- Any Linux/Windows Workstation with NVIDIA GPU (Docker):
# Lean public image (downloads the public weights on first boot); tags and digests: docs/deploy/public-images.md docker run --gpus all -p 8080:8080 us-central1-docker.pkg.dev/dgem-diffusiongemma/dgem/dgem:v0.1.0@sha256:5fa4a866163169aaf91c3bdb727020ff3c86e26e84ffadad86bc859053d45b26 - Remote Google Cloud Endpoint:
export DGEM_VERTEX_URL="<endpoint-id>" GCP_PROJECT_NUMBER="<project-number>" DGEM_GCP_AUTH=1
make build # compiles ./bin/dgem
./bin/dgem decide -t templates/support_triage.json.tmpl \
-v 'ticket=I was billed $500 twice for my annual renewal this morning!' --stats./bin/dgem serve --port 8090Open http://localhost:8090 to inspect all 26+ templates, evaluate Stage 2 Gemini cascades, and visualize OpenTelemetry trace waterfalls.
Next, pick your path: build, deploy and operate Β· confidence and calibration Β· write your first policy.
In one sentence: IDC makes a
dgemconfidence score reflect the question, not the position where each answer was listed, and flags decisions whose answer depends on the list order. Plain-English guide with a worked example:docs/confidence-beyond-shannon.md.
The problem. Per-question Shannon entropy (
What IDC does.
-
Null-Prior De-Biasing (
--null-prior-debias, no labeled data): divides out the measured slot habit. On the 50-item calibration suite it improved Brier from 0.175β0.193 (three same-session baselines) to 0.147; on the 231-item JevBench set it did not help (186 vs 187 correct, worse calibration). Suite-dependent, so validate before enabling. -
Dual-Mirror Canvas (
--dual-mirror): adds a reversed-order copy of each choice question to the same canvas, so the forward and reversed readings come from one forward pass, and their gap (Mirror TVD) flags order-dependent answers. A research diagnostic: with lettered options the reversed slot lowers forward accuracy (letter collision, EXP-15); digit-labelled mirrors (--mirror-mode reversed-digits) stay within noise but add little error detection beyond hesitation (EXP-15βEXP-17). Not recommended in production. -
Slot Temperature Scaling (
EXP-11): softens over-sharp scores. Needs labeled data. Fitted on held-out folds it cut ECE by 24β33% on 231 JevBench items ($T^* \approx 1.5$ ) but gave no reliable gain on the 50-item suite.
What's new. Removing a content-free prior (contextual calibration, Zhao et al. 2021), permutation debiasing (e.g. PriDe, Zheng et al. 2023), and temperature scaling (Guo et al. 2017) are known techniques. The part specific to a diffusion decision model is checking a reversed ballot on every request without a second forward pass, which turns order sensitivity from an offline audit into a per-request signal.
Status. A same-session re-run on 50 + 231 items (EXP-14, versioned receipts in benchmarks/runs/) and follow-ups (EXP-15βEXP-17) gave mixed results: the order-bias problem is real and reproducible, but the corrections are suite-dependent, and hesitation gating remains the recommended production signal. IDC is CLI-only today. See IDC Β§6 for every number and Proposed Experiments for what comes next.
| Result | Value | Sample | Receipt |
|---|---|---|---|
Single-pass accuracy, 11 public datasets (dgem bench-calibration) |
88.0% (44/50) | 50 | benchmarks/results_calibration_cloudrun.json |
| + Null-prior de-biasing (IDC), same session | 45/50, Brier 0.147 vs 0.175β0.193 (3 baselines) | 50 | benchmarks/runs/20260925-vertex-idc/ |
+ Entropy cascade to gemini-3.8-flash ( |
98.0% (49/50), 56% lower cost than Gemini on every item | 50 | results_calibration_cascade_normalized.json |
JevBench v1.3.1: dgem single pass β entropy cascade (hesitation β₯ 16%, 39% escalated, offline) |
187β189 β 221 of 231 | 231 | benchmarks/runs/20260925-vertex-idc/ |
| JevBench v1.3.1: null-prior de-biasing | 186 of 231, Brier 0.293 vs 0.264 (no gain) | 231 | benchmarks/runs/20260925-vertex-idc/ |
| Decision Index panel: bracket routing (> 26 options) + slot batching | 76.67 β 98.89, coverage 16/22 β 22/22 | 22 requests | benchmarks/decision_index/ |
Listwise reranking of 10 passages in one pass (EXP-10) |
0.9265 nDCG@10, 0% ties | 30 queries | results_rerank_cloudrun.json |
Content-free Slot-A habit (EXP-13B) |
88.3% / 78.3% / 49.3% for K = 2 / 3 / 4 | probe | results_permutation_cloudrun.json |
Run-to-run noise is about Β±1 item on 50 and Β±2 on 231. Cascade thresholds were chosen on the evaluation items; treat these as directional. Compare runs with python3 scripts/bench_runs.py compare. Details and caveats: Benchmark Report, Experiment Ledger.
See Decision Studio, MCP and HTTP API and Gateway and routing for full details:
| Interaction Surface | Command / Endpoint | Description |
|---|---|---|
| 1. π₯οΈ Decision Studio Web App | ./bin/dgem serve --port 8090 |
Embedded Lit WebComponents web application featuring all 26+ .json.tmpl decision policies (core, calibration, multimodal, rerank), topbar Backend Target selector (vertex_first | vertex | cloudrun), Stage 2 Gemini Cascade (gemini-3.8-flash), live SigLIP 2D Bounding Box SVG overlays (EXP-09), a plain-English Concepts tab (including an IDC walkthrough), and OpenTelemetry Trace Waterfall inspection. |
2. π€ Model Context Protocol (MCP) |
./bin/dgem mcp (stdio)POST /mcp (Streamable HTTP) |
Native MCP server exposing 6 tools (decide_policy, locate_bounding_boxes, decide_custom_questions, list_policy_templates, get_health_and_gpu_status, warmup_gpu) with backend (vertex_first | vertex | cloudrun) and Stage 2 Gemini Cascade support (cascade_mode, cascade_threshold, cascade_model). |
| 3. π HTTP Gateway REST API | POST /api/decide/{template}POST /v1/systemone, GET /api/templates |
Execute any .json.tmpl decision policy or /v1/systemone schema with X-DGem-Backend: vertex_first | vertex | cloudrun (X-DGem-Backend-Used returned on every response) and optional Stage 2 gemini-3.8-flash cascade. |
| 4. β¨οΈ CLI & 9 Benchmark Harnesses | ./bin/dgem decide --vertex-url ..../bin/dgem bench-* |
Direct single-pass decisions (--stats, --null-prior-debias, --dual-mirror) and nine reproducible evaluation harnesses (bench, bench-ecotone, bench-intents, bench-calibration, bench-bbox, bench-rerank, bench-jev, bench-decision-index, bench-permutation) backed by docs/experiments/ (EXP-01 β EXP-13). |
Production teams have historically chosen between two extremes for automated triage, routing, and guardrails:
- Discriminative classifiers & automata (BERT / DeBERTa / C++ WFSTs): very fast, but rigid. A new policy rule or category means new labeled data, retraining, and redeployment.
- Autoregressive LLMs (Gemini / GPT / Gemma 4): zero-shot flexible, but they generate answers token by token (seconds per multi-field JSON answer), can drift from the output format, and their token probabilities are spread across formatting tokens rather than the decision itself.
A zero-shot decision model sits in between. dgem compiles a .json.tmpl template into a fixed diffusion canvas (32β256 tokens) with bidirectional attention; boolean gates, [AβZ] choices, and ordinal scores are read together in one forward pass, and every answer is constrained to the allowed options.
| Architectural Dimension | Discrete Diffusion Decision Model (dgem) |
Discriminative Encoder (DeBERTa-v3 / Llama-Guard) | Autoregressive LLM (Gemini / Gemma 4) | Compiled Rulebook (ecotone C++ WFST) |
|---|---|---|---|---|
| Policy Adaptability |
Zero-shot Policy-as-Template (edit .json.tmpl) |
Labeled dataset & retraining per label change | Zero-shot prompt engineering | Manual grammar authoring & compilation |
| Inference Latency | ~55 ms GPU / ~150 ms end to end for a 3-question decision on RTX PRO 6000; ~1.4 s for a 12-slot rerank (measured on L4) | ~5 β 25 ms (single head) | 17,486.6 ms (~17.5 s for 3-slot JSON + CoT) |
1.35 β 8.68 ms (1.54 ms p50 over UDS) |
| Passes per Request | 1 forward pass for all questions (cost grows with canvas length) | One classifier per attribute | One token per step ($O(T_{\text{output}})$) |
|
| Joint Slot Conditioning |
Bidirectional (slot_1 <-> slot_2) in a single pass |
Independent heads | Left-to-right only | Local sliding window (1β3 tokens) |
| Uncertainty & Calibration | Per-option probabilities + entropy; label-free order-bias correction (null-prior) and same-pass reversed-ballot check (IDC) | Often overconfident out-of-distribution | Sequence-level logprobs over formatting tokens | Static arc weights |
| Guardrail Examples (50-item suite) |
AgentDrift 7/7, prompt injection 4/4, RAG grounding 2/2 |
Narrow single-task scope | High accuracy, 15β25Γ higher latency | 36.7% on semiotic polysemy traps |
Measured with the current serving image (3-question decision, samples: 1, p50;
receipts):
| Stage | Where | Cold start | GPU / end to end | Guide |
|---|---|---|---|---|
| Crawl | Apple Silicon (diffgemma) or a local NVIDIA GPU |
None | ~0.9 s on Metal | Run on your laptop |
| Walk | Cloud Run GPU, 1Γ RTX PRO 6000, scale to zero | 2.3β2.7 min | 61 / 153 ms | Deploy on Cloud Run |
| Run | Vertex AI dedicated endpoint, g4-standard-48 + 1Γ RTX PRO 6000 |
None | 55 / 150 ms | Production on Vertex AI |
A gateway (dgem serve) in front routes to Vertex first and fails over to Cloud Run
(Gateway and routing). Capacity, runbook and observability:
Latency and capacity Β· Operations runbook.
# Clone the repository
git clone https://github.com/ghchinoy/dgem.git
cd dgem
# Compile dgem binary into bin/
make buildEvaluate customer tickets, code changes, or security alerts in a single sub-second forward pass:
./bin/dgem decide -t templates/support_triage.json.tmpl \
-v 'ticket=I was billed $500 twice for my annual renewal this morning!' \
--statsOutput:
QUESTION | TYPE | VALUE / CHOICE | CONFIDENCE | ENTROPY (H) | AGREEMENT
-----------------------------------------------------------------------------------------
sentiment | score | frustrated | 99.8% | 0.002 nats | 1.00
team | choice | billing | 100.0% | 0.000 nats | 1.00
urgent | boolean | yes | 99.9% | 0.001 nats | 1.00
ββββββββββββββββββββββββββββββββ STATS ββββββββββββββββββββββββββββββββ
Model: nvidia/diffusiongemma-26B-A4B-it-NVFP4
Endpoint: http://127.0.0.1:8080/v1/chat/completions
Total Wall Time: 856 ms
KV Cache Reused: 169 tokens (82.8% hit rate)
Denoise Steps: 1 step (policy: samples=4)
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Standard chat completion with optional thinking mode:
./bin/dgem ask "Explain discrete block diffusion in two sentences."Connect to any remote GCE or Cloud Run GPU service:
./bin/dgem decide \
-u "http://<EXTERNAL_IP>:8080/v1" \
-m "nvidia/diffusiongemma-26B-A4B-it-NVFP4" \
-t templates/support_triage.json.tmpl \
-v 'ticket=Outage: production database cluster unreachable' \
--statsAttach local image paths (automatically base64 encoded) or remote URLs:
./bin/dgem decide -t templates/multimodal/ui_design_review.json.tmpl \
-I fixtures/ui_component.svg \
-v 'component=CheckoutCard' \
--statsdgem includes nine benchmark harnesses (all tracked in docs/experiments/README.md). Four of the most commonly used are below; the others are bench-jev (JevBench v1.3.1), bench-decision-index (Decision Index panel + /v1/systemone), bench-permutation (option-order sensitivity and IDC, EXP-13), bench-rerank (listwise reranking, EXP-10), and bench-bbox (bounding boxes, EXP-09).
Evaluates 50 items across 11 public datasets (benchmarks/calibration_suite.jsonl), testing declarative policy templates (templates/calibration/*.json.tmpl) across agent trajectory hijacking (AgentDrift), multilingual jailbreaks (deepset/prompt-injections), RAG fact grounding (LLM-AggreFact), retrieval relevance (MS MARCO), toxicity (Jigsaw Civil Comments), and human annotator disagreement (ChaosNLI):
./bin/dgem bench-calibration -u "https://<CLOUD_RUN_URL>/v1" --gcp-auth -w 4 \
-o benchmarks/results_calibration_cloudrun.json| Public Dataset / Policy Domain | Cases | Accuracy | Mean |
Mean Entropy |
Avg Latency |
|---|---|---|---|---|---|
AgentDrift (agent_step_drift.json.tmpl β Hijack + 4-Way Step Localization) |
7 | 100.0% (7/7) β | 0.997 |
0.0186 nats |
693 ms |
deepset/prompt-injections (prompt_injection.json.tmpl β en/de Gate) |
4 | 100.0% (4/4) β | 0.980 |
0.0817 nats |
669 ms |
LLM-AggreFact & MS MARCO (RAG Grounding & Retrieval Relevance) |
4 | 100.0% (4/4) β | 0.993 |
0.0403 nats |
728 ms |
CLINC150, Banking77, GoEmotions, BoolQ, Yelp/SST-5 |
20 | 100.0% (20/20) β | 0.898 |
0.3263 nats |
769 ms |
ChaosNLI Crowd Consensus (low-entropy) |
3 | 100.0% (3/3) | 0.986 |
0.0744 nats (1.0Γ) |
625 ms |
ChaosNLI Crowd Split (high-entropy) |
3 | 33.3% (1/3) | 0.759 |
0.5932 nats (8.0Γ higher; n=3) |
731 ms |
Stage 1 Alone: DiffusionGemma (steps=1, think=0) |
50 | 88.0% (44/50) | 0.925 |
0.2279 nats |
712 ms |
Raw Entropy Cascade (EXP-05a): dgemma [H<0.35] gemini-3.8-flash |
50 | 94.0% (47/50, +6.0%) |
0.959 |
0.1410 nats |
1,824 ms (72% early-exit) |
Normalized + Prior-Guided Cascade (EXP-05b, |
50 |
98.0% (49/50, +10.0%) β |
0.960 |
0.1416 nats ( |
2,105 ms (66% early-exit) |
Stage 2 Alone: gemini-3.8-flash (100% Frontier LLM) |
50 | 98.0% (49/50) | 0.959 |
0.1347 nats |
3,412 ms (4.8Γ slower) |
Evaluates 30 multi-field test cases (boolean + choice + score in a single pass) across support, code_review, and security (benchmarks/eval_dataset.jsonl):
./bin/dgem bench -d benchmarks/eval_dataset.jsonl -M slot -o benchmarks/results_cloudrun.jsonEvaluates 49 Text Normalization cases comparing C++ ecotone (OpenFst / Sparrowhawk WFSTs over unix:///tmp/ecotone.sock) against DiffusionGemma across semiotic polysemy traps and deterministic NSWs:
./bin/dgem bench-ecotone -c benchmarks/ecotone/tn_semiotics.jsonl --samples 1 -o benchmarks/results_ecotone.jsonEvaluates 30-way to 151-way intent routing and Out-of-Scope (oos) rejection on PolyAI/banking77 and DeepPavlov/clinc150:
./bin/dgem bench-intents --dataset banking77 --full --workers 16
./bin/dgem bench-intents --dataset clinc150 --full --workers 16Published at ghchinoy.github.io/dgem (index), organized by audience:
- Build, deploy and operate: From laptop to production Β· Run on your laptop Β· Use a remote GPU Β· Deploy on Cloud Run Β· Production on Vertex AI Β· Gateway and routing Β· Latency and capacity Β· Operations runbook Β· Observability
- Confidence and calibration: Overview Β· Calibrate your policy Β· Confidence beyond Shannon (IDC) Β· The journey to decision models Β· Glossary Β· Benchmark report Β· Experiment ledger Β· Proposed experiments
- Policies and decisions: Your first decision policy Β· Authoring and Stage 2 cascades Β· Run a dataset Β· Template catalog Β· Real-world applications Β· Taxonomy discovery
- Reference: CLI, HTTP and MCP Β· Studio, MCP and HTTP API Β· Public container images Β· Vertex AI vs. Cloud Run Β· Apple Silicon engine Β· How the model decides in one pass Β· Ecotone comparison Β· Engineering history (archive)
Issues, bug reports, and feature discussions are welcome! However, we are not accepting pull requests (PRs) at this time. If you encounter a bug or have feedback on benchmark methodologies or templates, please open an Issue.
This project is licensed under the Apache-2.0 License.
Caution
This is not an officially supported Google product. This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.