| title | Decision Model Experiment Ledger (EXP-01 – EXP-17) |
|---|---|
| description | Structured empirical research log tracking DiffusionGemma as a Zero-Shot Decision Model, Declarative Policy-as-Template evaluation, Epistemic Calibration, Listwise Reranking, JevBench v1.3.1 Parity, Decision Index, Permutation Invariance, and Next-Horizon Cascades. |
This directory serves as the structured research and engineering log for DiffusionGemma (dgemma) as a Zero-Shot Decision Model and dgem as a Declarative Policy Engine (Policy-as-Template).
Much like benchmarks/ stores reproducible JSONL evaluation suites and JSON telemetry receipts, docs/experiments/ chronicles why each experiment was designed, how its templates encode domain policy, what the empirical single-pass logprobs and Shannon entropy
📋 Planned work: Experiments we intend to run, each with a pre-registered hypothesis, design, and decision criteria, are tracked in the Proposed Experiments Register (
PROP-00–PROP-10). When an entry starts running, it gets the nextEXP-XXnumber here.
Every experiment in dgem connects three version-controlled artifacts:
flowchart LR
T["1. Policy Templates\n(templates/**/*.json.tmpl)"] --> H["2. CLI Harness\n(cmd/dgem/bench*.go)"]
D["2. JSONL Datasets\n(benchmarks/*.jsonl)"] --> H
H --> R["3. Telemetry Receipts\n(benchmarks/results_*.json)"]
R --> E["4. Experiment Log\n(docs/experiments/*.md)"]
| Directory / Path | Purpose | Format |
|---|---|---|
templates/ |
Executable Decision Policies (choice & score slots, depends_on / ask_if DAGs) |
.json.tmpl |
templates/calibration/ |
Public Dataset Calibration & Guardrail Policies (AgentDrift, ChaosNLI, LLM-AggreFact, prompt-injections) |
.json.tmpl |
templates/rerank/ |
Listwise & Pointwise Neural Reranking + RAG Security Policies (EXP-10) |
.json.tmpl |
benchmarks/*.jsonl |
Evaluation Datasets (eval_dataset.jsonl, calibration_suite.jsonl, rerank_suite.jsonl, jevbench/jevbench_public.jsonl, banking77_26.jsonl, clinc150_26.jsonl, tn_*.jsonl) |
.jsonl |
benchmarks/results_*.json |
Immutable Telemetry Receipts (logprobs, slot probabilities |
.json |
docs/experiments/ |
Experiment Ledger & Architectural Deep-Dives | .md |
| ID | Experiment Title | Policy Templates | Dataset (benchmarks/) |
CLI Command | Primary Finding & Telemetry Receipt | Status |
|---|---|---|---|---|---|---|
EXP-01 |
Multi-Domain Joint Slot Readout Across 4 Hardware Tiers |
support_triage.json.tmplcode_review.json.tmplsecurity_audit.json.tmpl
|
eval_dataset.jsonl (30 cases, 113 slots) |
dgem bench |
93.3% accuracy (bfloat16 2× A100), 86.7% (4-bit Metal), 80.0% at 458.9 ms (Serverless Cloud Run 1× L4). Zero JSON syntax errors across all runs.Receipts: results_local_metal_slot.json, results_gce_a100_16bit.json, results_cloudrun.json
|
✅ Completed |
EXP-02 |
Contextual Text Normalization: WFST (ecotone) vs. DiffusionGemma |
tn_semiotics.json.tmpltn_audit.json.tmpl
|
tn_semiotics.jsonl (30 polysemy traps)tn_challenge_en.jsonl (19 NSWs) |
dgem bench-ecotone |
96.7% accuracy on contextual homographs (St. 1/2 Receipt: results_ecotone_comparison.json
|
✅ Completed |
EXP-03 |
High-Cardinality Intent Routing & Out-of-Scope (oos) Detection |
banking77.json.tmplclinc150.json.tmpl(26-option [A-Z] slice) |
banking77_26.jsonl (100 items)clinc150_26.jsonl (100 items, 16% oos) |
dgem bench-intents |
92.0% accuracy on PolyAI/banking77 (H = 0.2518 nats, 485 ms) and 95.0% accuracy on DeepPavlov/clinc150 (H = 0.1459 nats, 562 ms) with 93.8% zero-shot oos recall.Receipts: results_intents_banking77_cloudrun.json, results_intents_clinc150_cloudrun.json
|
✅ Completed |
EXP-04 |
Public Dataset Calibration, Guardrails & ChaosNLI Epistemic Entropy |
templates/calibration/*.json.tmpl (8 policy templates) |
calibration_suite.jsonl (50 items across 11 public datasets) |
dgem bench-calibration |
88.0% overall accuracy (44/50) at 712 ms. 100% accuracy on AgentDrift (7/7), prompt-injections (4/4), LLM-AggreFact (2/2), and MS MARCO (2/2). On ChaosNLI, slot entropy 0.0744 nats 0.5932 nats) with human annotator disagreement.Receipt: results_calibration_cloudrun.json
|
✅ Completed |
EXP-05 |
Entropy-Gated Escalation Cascade ( |
templates/calibration/*.json.tmpl |
calibration_suite.jsonl (50 items) |
dgem bench-calibration --cascade-from ... --normalize-entropy --cascade-threshold 0.16 |
98.0% cascade accuracy (49/50, +10.0% gain, 100% on ANLI-R3 & 100% on Adversarial + Ambiguous tiers) using Cardinality-Normalized Entropy (gemini-3.8-flash (98.0%) while saving 66% of frontier LLM calls.Receipts: results_calibration_cascade_normalized.json, results_calibration_cascade_prior_guided.json, results_calibration_cascade.json
|
✅ Completed (Analysis) |
Detailed architectural specifications, mathematical formulations, and empirical cascade results for EXP-05 through EXP-08 are documented in exp-05-roadmap-cascades-and-dags.md, EXP-10 is documented in exp-10-listwise-diffusion-reranking.md, EXP-11 is documented in exp-11-jevbench-parity.md, EXP-12 is documented in exp-12-decision-index.md, and EXP-13 is documented in exp-13-permutation-invariance.md.
| ID | Experiment Title | Core Hypothesis | Target Datasets & Templates | Target CLI Flag / Feature | Status |
|---|---|---|---|---|---|
EXP-06 |
Decision Models vs. Discriminative Encoder Heads | Compare zero-shot dgemma (Policy-as-Template) against fine-tuned encoder heads (DeBERTa-v3-large, Llama-Guard-3-8B, ModernBERT) across joint multi-slot capability, policy adaptability (0s template edit vs. fine-tuning), and Expected Calibration Error (ECE). |
calibration_suite.jsonl (AgentDrift, prompt-injections, ChaosNLI) |
dgem bench-encoders |
🔬 Planned (Spec) |
EXP-07 |
Conditional Policy DAGs (depends_on & ask_if) |
Multi-stage conditional templates prune irrelevant downstream branches when upstream gate slots resolve negative, cutting slot density and eliminating contradictory sub-slot classifications. | templates/secops_conditional_dag.json.tmpl |
Native structured_server.py DAG execution (depends_on, ask_if) |
🧪 Template Ready (Spec) |
EXP-08 |
Multimodal Vision & Document Policy Readout (SigLIP) |
Because DiffusionGemma inherits Gemma 4's SigLIP vision encoder (896×896 patches), bidirectional slot readout can classify receipts, UI screenshots, and PDF invoices in a single forward pass (~500 ms). |
Receipt & UI compliance image suite | dgem decide --image <path> |
🔬 Planned (Spec) |
EXP-09 |
Single-Pass Spatial Grounding, Softmax-Expectation Sub-Bin Regression & Per-Edge Occlusion Entropy | Predicting a 2D bounding box [ymin, xmin, ymax, xmax] as 4 parallel 21-bin (00..100) slots in think=0 (reads=1): (1) On live Cloud Run dgemma (SigLIP enabled), Softmax Expectation ($\hat{c}m = \sum_k v_k p{m,k}$) improves mIoU from 0.2898 to 0.3773 (+8.75% absolute / +30.2% relative, Acc@0.5 0% 18.2%; up to +50.4% IoU on narrow stemware in 008.png and 0.9866 vs 0.7997 simulated), (2) Per-edge normalized entropy (1.37× on occluded box edges (0.6810 vs 0.4970 live; 2.86× simulated), and (3) DETR-style parallel object query slots prevent duplicate collapse via bidirectional self-attention. |
templates/multimodal/bbox_localization.json.tmpltemplates/multimodal/bbox_multi_object_detr.json.tmpltemplates/multimodal/bbox_multi_object_set.json.tmplbenchmarks/bbox_suite.jsonl
|
dgem bench-bbox --annotateresults_bbox_cloudrun.jsonresults_bbox_simulated.json
|
✅ Completed |
EXP-10 |
Listwise Diffusion Canvas Reranking, Softmax Expectation & RAG Poison Quarantine | Evaluate 10 candidate passages (doc_01..doc_10) + 2 RAG security/abstention gates (answer_present, poisoned_passage) simultaneously in 1 forward pass (12 slots, ~138 ms effective/passage on Cloud Run 1× L4). Continuous Softmax Expectation ($\hat{r}i = \sum{g=0}^3 g \cdot p_{i,g}$) cuts Exact Tie Rate from 70.0% (discrete argmax) to 0.0%, lifting nDCG@10 from 0.8416 to 0.9265 (+8.49 pts) and MRR@10 from 0.7407 to 0.9444 (+20.37 pts), with 100.0% NevIR negation accuracy, +0.7533 FollowIR p-MRR policy steerability, and 100.0% prompt-injection quarantine. |
templates/rerank/listwise_decision_rerank.json.tmpltemplates/rerank/pointwise_rerank.json.tmplbenchmarks/rerank_suite.jsonl
|
dgem bench-rerankresults_rerank_cloudrun.json
|
✅ Completed (Analysis) |
EXP-11 |
JevBench v1.3.1 & v1.4 4-Axis Parity, Harmonic Scoring, Slot Temperature Calibration & Upstream Sync |
Upgrade dgem with JevBench v1.3.1 & v1.4's 4-Axis Scorecards (Chance-Corrected Intelligence, 10-Bin ECE + Soft TVD Calibration, Speed with 2×+0.15s load factor adjustment, Cost, and v1.4 Harmonic Mean Composite with 3 gates EXP-04, T* = 1.35 cuts 10-bin ECE by 56.2% (0.0745 0.0326), lifting Calibration from 82.67 to 88.18 (76.79 Composite). On EXP-05, the Entropy-Gated Cascade scores 70.16 Composite (98.0% raw / 97.17% chance-corrected) vs. standalone gemini-3.8-flash at 62.76 (56% lower cost). Syncs and replays all 231 JevBench public tasks (75.70 v1.3.1 Composite #1, surpassing #1 Hopper [75.40]; 71.95* v1.4 Harmonic Estimate). Supports --compare-leaderboard against official v1.4.2 rows (Plumb-4B, decider-4b v2, djev, OpenJev). |
templates/jevbench_generic.json.tmplbenchmarks/jevbench/jevbench_public.jsonlbenchmarks/jevbench/manifest.lock.jsonbenchmarks/jevbench/leaderboard_v14.json
|
dgem bench-jev --sync --check-upstreamdgem bench-jev --compare-leaderboardresults_djev_upstream_calibrated.json
|
✅ Completed (Analysis) |
EXP-12 |
jev-decision-index (apolinario/decision-index) 5-Area Benchmark, Wide-Canvas Adapter & /v1/systemone Protocol |
Evaluate dgemma across the 5-Area Decision Index (19 scored panel benchmarks + 3 wide-option benchmarks). Eliminates HTTP 422 capacity rejections (0.0 score penalty on M > 10 slots in ContractNLI/BRIGHT/ToolRet and K > 26 options in API-Bank/BANKING77/CLINC150) via Multi-Slot Canvas Batching + 2-Stage Bracket Tournament Routing (T*=1.25), lifting structural coverage from 72.7% to 100.0% and Headline Decision Index from 76.67 to 98.89 (+22.22 pts, Skill 98.33, ECE 0.0371). Projects to Rank #3 / 49 Overall (#1 Diffusion, 46.90 balanced_skill) as Pure System-1 (0% LLM) and Rank #1 / 49 Overall (53.94 balanced_skill, beating TypeSafe Jev 1.13.0 [51.67]) with EXP-05/11 Entropy-Gated Cascades (#1 / 31 on v0.1 at 56.51 raw). |
benchmarks/decision_index/panel_suite.jsonlpkg/decisionindex/engine.go
|
dgem bench-decision-index --compare-naivedgem bench-decision-index --serve-systemone :8095results_decision_index_cloudrun.json
|
✅ Completed (Analysis) |
EXP-13 |
Permutation Sensitivity, Content-Free Null-Prior De-Biasing & |
Quantifies how option order distorts single-pass confidence. Content-free probes show a strong Slot-'A' preference (13B) cuts Brier 0.0173 0.0017 but raises cyclic flip rate 12.5% 25%; (2) Cyclic JSD (13A) is ~156× higher on the 4 ambiguous items, where reordering flips 2/4 answers; (3) Dual-Mirror (13C) reads forward + reversed slots in one pass (124.7 vs 128.8 ms) and flags perm_08 (99.9%, Mirror TVD = 0.258) but misses perm_06 (TVD 0.0006) and worsens Brier (0.0410). See the plain-English summary in Confidence Beyond Shannon (IDC). |
benchmarks/permutation_suite.jsonlpkg/permutation/permutation.go
|
dgem bench-permutationresults_permutation_cloudrun.json
|
✅ Completed (Analysis) |
EXP-14 |
Same-Session IDC Re-run on Vertex AI (versioned runs) | Pre-registered follow-ups PROP-01/02/05/10 plus an EXP-13 reproduction, run in one session with versioned receipts. Findings: ±1–2 item noise floor; null-prior helps on the 50-item suite (Brier 0.147 vs 0.175–0.193) but not on JevBench (186 vs 187, Brier worse); the dual-mirror slot id __mirror_rev degraded readings (JevBench 187 → 145; fixed to __rev), and a same-canvas mirror still lowers the forward reading; held-out temperature scaling cuts ECE 24–33% on 231 JevBench items but not on 50 items; entropy cascade reaches 221/231 at 39% escalation. Details
|
benchmarks/calibration_suite.jsonl, benchmarks/jevbench/jevbench_public.jsonl, benchmarks/permutation_suite.jsonl
|
scripts/run_idc_rerun.shscripts/analyze_idc.pybenchmarks/runs/20260925-vertex-idc*
|
✅ Completed |
EXP-15 |
Letter Collision in the Dual-Mirror (PROP-11) |
Same-canvas mirror damage comes from shared letters: on JevBench (231, Vertex G4, 3 baselines 182–189), an identical copy (184) and a digit-labelled reversed slot (182) are within noise, while the letter-labelled reversed slot drops to 154; forward answers change on 67% of same-letter items vs 6% otherwise (p≈1e-19). Details | benchmarks/jevbench/jevbench_public.jsonl |
dgem bench-jev --dual-mirror --mirror-modescripts/analyze_idc.py collisionbenchmarks/runs/20260925-prop11-letter-collision
|
✅ Completed |
EXP-16 |
Slot Names Are Part of the Prompt (PROP-16) |
Single-slot ids (q1, random, mirror, check) are within the baseline band (183–189) on JevBench; a loaded second-slot id (__mirror_rev, identical options so no letter collision) drops accuracy to 162 (second slot 135) vs 181 for __rev. Style rule: neutral ids in multi-question schemas. Details
|
benchmarks/jevbench/jevbench_public.jsonl |
dgem bench-jev --slot-idbenchmarks/runs/20260926-prop16-slot-names
|
✅ Completed |
EXP-17 |
Separate-Pass Mirror (PROP-12) |
Forward vs. separately-run reversed pass on JevBench (231, Vertex G4): order disagreement relates to errors beyond hesitation (mean partial Spearman 0.133, CI [0.025, 0.240]), but two forward passes show a similar non-significant value (0.083) and CV AUROC improves by ≤ 0.016 at double the cost. Details | benchmarks/jevbench/jevbench_public.jsonl |
dgem bench-jev --flip-optionsscripts/analyze_idc.py separate-passbenchmarks/runs/20260926-prop12-separate-pass
|
✅ Completed |
EXP-18 |
Judge Capabilities With and Without an Autoregressive Autorater (mizan Experiment 07) | Gold-scored comparison over 13 human-labelled judge suites (n=100 each) and 650 computation items, rerun on serving image v0.1.0 (2026-09-29) and replicated from 2026-09-26: v0.1.0 shows no regression (67.0 → 67.9%, p=0.23). DiffusionGemma matches gemini-3.8-flash on safety (80 vs 76), faithfulness (81 vs 78) and RewardBench with a mirror read (90 vs 90); trails on toxicity (81 vs 87.9, p=0.065) and LLMBar instruction following (79 vs 94, p=0.0015); MT-Bench expert preference favours Gemini by 5–10 points on both dates. A fixed 35%-hesitation cascade reaches 81/86 on safety/faithfulness with 8–32% escalation; p50 105 ms (G4) vs 3,125 ms. No-model metrics: mizan local = Vertex on exact/tool/trajectory (350/350). Speed/cost: one G4 replica ~83 items/s (0 errors sustained); ~$0.02 per 1k judgements fully utilized vs $1.06–2.25 (3.8-flash) and ≥$10 (Vertex predefined metrics); live cascade matches the simulation. Details | mizan judge-eval pack; suites built by mizan scripts/judge-eval/build_suites.py (BeaverTails, ToxicChat, HaluBench, MT-Bench human, RewardBench, HelpSteer2, SummEval) |
mizan eval compare-enginesmizan scripts/judge-eval/run_all.sh
|
✅ Completed |
Across 280+ evaluated cases (EXP-01 through EXP-04), DiffusionGemma demonstrates three properties that distinguish a Decision Model from both autoregressive generative LLMs and discriminative encoder classifiers:
-
Zero-Shot Policy-as-Template (
0sPolicy Iteration): Adding a new governance domain (AgentDriftbehavioral drift,LLM-AggreFactRAG grounding,deepset/prompt-injections) required zero gradient updates, zero labeled training splits, and zero output regex parsers. Writing a.json.tmplfile immediately turned the 9B model into a 100%-accurate classifier across those four benchmarks (15/15combined inEXP-04). -
Joint Multi-Slot Co-Adaptation in
$O(1)$ Forward Pass: InEXP-01(support_triage,code_review,security_audit) andEXP-04(AgentDrift), a single forward pass (458.9 mson 1× L4) evaluates 3 to 5 orthogonal decision dimensions simultaneously (drift_detected+drift_severity+remediation_action). Because the[MASK]canvas is bidirectional, the severity and remediation slots attend directly to the detection slot in the same pass. -
Intrinsic Epistemic Calibration (
EXP-04ChaosNLI8.0× Entropy Multiplier): Traditional neural classifiers suffer from overconfidence on out-of-distribution or genuinely ambiguous inputs. InEXP-04, evaluatingChaosNLI(where 100 human annotators rated each premise/hypothesis pair) proved that DiffusionGemma's restricted-softmax Shannon entropy$H$ scales monotonically with human disagreement:-
High human consensus (
low-entropy): 100.0% accuracy,$H = 0.0744\text{ nats}$ (1.0×baseline) -
Moderate split (
ambiguous):$H = 0.4061\text{ nats}$ (5.5×entropy increase) -
Near-uniform 3-way human split (
high-entropy):$H = 0.5932\text{ nats}$ (8.0×entropy increase) -
Adversarial multi-hop reasoning (
ANLI-R3): When single-passthink: 0readout fails (0/3), slot entropy automatically spikes to$H = 0.4104\text{ nats}$ (5.5×), providing an unmistakable mathematical trigger ($H > 0.30\text{ nats}$ ) to escalate to a reasoning pass (EXP-05).
-
High human consensus (