Skip to content

Latest commit

 

History

History
89 lines (68 loc) · 21.5 KB

File metadata and controls

89 lines (68 loc) · 21.5 KB
title Decision Model Experiment Ledger (EXP-01 – EXP-17)
description Structured empirical research log tracking DiffusionGemma as a Zero-Shot Decision Model, Declarative Policy-as-Template evaluation, Epistemic Calibration, Listwise Reranking, JevBench v1.3.1 Parity, Decision Index, Permutation Invariance, and Next-Horizon Cascades.

DiffusionGemma (dgem) Experiment Ledger

This directory serves as the structured research and engineering log for DiffusionGemma (dgemma) as a Zero-Shot Decision Model and dgem as a Declarative Policy Engine (Policy-as-Template).

Much like benchmarks/ stores reproducible JSONL evaluation suites and JSON telemetry receipts, docs/experiments/ chronicles why each experiment was designed, how its templates encode domain policy, what the empirical single-pass logprobs and Shannon entropy $H$ revealed, and where our next architectural frontiers lie.

📋 Planned work: Experiments we intend to run, each with a pre-registered hypothesis, design, and decision criteria, are tracked in the Proposed Experiments Register (PROP-00 – PROP-10). When an entry starts running, it gets the next EXP-XX number here.


1. Repository Taxonomy & Filepath Architecture

Every experiment in dgem connects three version-controlled artifacts:

flowchart LR
    T["1. Policy Templates\n(templates/**/*.json.tmpl)"] --> H["2. CLI Harness\n(cmd/dgem/bench*.go)"]
    D["2. JSONL Datasets\n(benchmarks/*.jsonl)"] --> H
    H --> R["3. Telemetry Receipts\n(benchmarks/results_*.json)"]
    R --> E["4. Experiment Log\n(docs/experiments/*.md)"]
Loading
Directory / Path Purpose Format
templates/ Executable Decision Policies (choice & score slots, depends_on / ask_if DAGs) .json.tmpl
templates/calibration/ Public Dataset Calibration & Guardrail Policies (AgentDrift, ChaosNLI, LLM-AggreFact, prompt-injections) .json.tmpl
templates/rerank/ Listwise & Pointwise Neural Reranking + RAG Security Policies (EXP-10) .json.tmpl
benchmarks/*.jsonl Evaluation Datasets (eval_dataset.jsonl, calibration_suite.jsonl, rerank_suite.jsonl, jevbench/jevbench_public.jsonl, banking77_26.jsonl, clinc150_26.jsonl, tn_*.jsonl) .jsonl
benchmarks/results_*.json Immutable Telemetry Receipts (logprobs, slot probabilities $p_i$, Shannon entropy $H$, wall latency) .json
docs/experiments/ Experiment Ledger & Architectural Deep-Dives .md

2. Master Experiment Index (EXP-01 – EXP-18)

Part A — Completed Empirical Studies

ID Experiment Title Policy Templates Dataset (benchmarks/) CLI Command Primary Finding & Telemetry Receipt Status
EXP-01 Multi-Domain Joint Slot Readout Across 4 Hardware Tiers support_triage.json.tmpl
code_review.json.tmpl
security_audit.json.tmpl
eval_dataset.jsonl (30 cases, 113 slots) dgem bench 93.3% accuracy (bfloat16 2× A100), 86.7% (4-bit Metal), 80.0% at 458.9 ms (Serverless Cloud Run 1× L4). Zero JSON syntax errors across all runs.
Receipts: results_local_metal_slot.json, results_gce_a100_16bit.json, results_cloudrun.json
✅ Completed
EXP-02 Contextual Text Normalization: WFST (ecotone) vs. DiffusionGemma tn_semiotics.json.tmpl
tn_audit.json.tmpl
tn_semiotics.jsonl (30 polysemy traps)
tn_challenge_en.jsonl (19 NSWs)
dgem bench-ecotone 96.7% accuracy on contextual homographs (St. $\rightarrow$ Saint vs Street, 1/2 $\rightarrow$ January second vs one half) where deterministic WFSTs score 50.0%, proving bidirectional context resolves semiotic ambiguity.
Receipt: results_ecotone_comparison.json
✅ Completed
EXP-03 High-Cardinality Intent Routing & Out-of-Scope (oos) Detection banking77.json.tmpl
clinc150.json.tmpl
(26-option [A-Z] slice)
banking77_26.jsonl (100 items)
clinc150_26.jsonl (100 items, 16% oos)
dgem bench-intents 92.0% accuracy on PolyAI/banking77 (H = 0.2518 nats, 485 ms) and 95.0% accuracy on DeepPavlov/clinc150 (H = 0.1459 nats, 562 ms) with 93.8% zero-shot oos recall.
Receipts: results_intents_banking77_cloudrun.json, results_intents_clinc150_cloudrun.json
✅ Completed
EXP-04 Public Dataset Calibration, Guardrails & ChaosNLI Epistemic Entropy templates/calibration/*.json.tmpl (8 policy templates) calibration_suite.jsonl (50 items across 11 public datasets) dgem bench-calibration 88.0% overall accuracy (44/50) at 712 ms. 100% accuracy on AgentDrift (7/7), prompt-injections (4/4), LLM-AggreFact (2/2), and MS MARCO (2/2). On ChaosNLI, slot entropy $H$ scales monotonically by 8.0× (0.0744 nats $\rightarrow$ 0.5932 nats) with human annotator disagreement.
Receipt: results_calibration_cloudrun.json
✅ Completed
EXP-05 Entropy-Gated Escalation Cascade ($\tilde{H}_m = H_m / \ln|\mathcal{V}_m|$ + Prior Forwarding) templates/calibration/*.json.tmpl calibration_suite.jsonl (50 items) dgem bench-calibration --cascade-from ... --normalize-entropy --cascade-threshold 0.16 98.0% cascade accuracy (49/50, +10.0% gain, 100% on ANLI-R3 & 100% on Adversarial + Ambiguous tiers) using Cardinality-Normalized Entropy ($\tilde{H} \ge 0.16$) and Pass-1 Slot Prior Forwarding, matching 100% standalone gemini-3.8-flash (98.0%) while saving 66% of frontier LLM calls.
Receipts: results_calibration_cascade_normalized.json, results_calibration_cascade_prior_guided.json, results_calibration_cascade.json
✅ Completed (Analysis)

Part B — Active & Next-Horizon Experiments (EXP-06 – EXP-18)

Detailed architectural specifications, mathematical formulations, and empirical cascade results for EXP-05 through EXP-08 are documented in exp-05-roadmap-cascades-and-dags.md, EXP-10 is documented in exp-10-listwise-diffusion-reranking.md, EXP-11 is documented in exp-11-jevbench-parity.md, EXP-12 is documented in exp-12-decision-index.md, and EXP-13 is documented in exp-13-permutation-invariance.md.

ID Experiment Title Core Hypothesis Target Datasets & Templates Target CLI Flag / Feature Status
EXP-06 Decision Models vs. Discriminative Encoder Heads Compare zero-shot dgemma (Policy-as-Template) against fine-tuned encoder heads (DeBERTa-v3-large, Llama-Guard-3-8B, ModernBERT) across joint multi-slot capability, policy adaptability (0s template edit vs. fine-tuning), and Expected Calibration Error (ECE). calibration_suite.jsonl (AgentDrift, prompt-injections, ChaosNLI) dgem bench-encoders 🔬 Planned (Spec)
EXP-07 Conditional Policy DAGs (depends_on & ask_if) Multi-stage conditional templates prune irrelevant downstream branches when upstream gate slots resolve negative, cutting slot density and eliminating contradictory sub-slot classifications. templates/secops_conditional_dag.json.tmpl Native structured_server.py DAG execution (depends_on, ask_if) 🧪 Template Ready (Spec)
EXP-08 Multimodal Vision & Document Policy Readout (SigLIP) Because DiffusionGemma inherits Gemma 4's SigLIP vision encoder (896×896 patches), bidirectional slot readout can classify receipts, UI screenshots, and PDF invoices in a single forward pass (~500 ms). Receipt & UI compliance image suite dgem decide --image <path> 🔬 Planned (Spec)
EXP-09 Single-Pass Spatial Grounding, Softmax-Expectation Sub-Bin Regression & Per-Edge Occlusion Entropy Predicting a 2D bounding box [ymin, xmin, ymax, xmax] as 4 parallel 21-bin (00..100) slots in think=0 (reads=1): (1) On live Cloud Run dgemma (SigLIP enabled), Softmax Expectation ($\hat{c}m = \sum_k v_k p{m,k}$) improves mIoU from 0.2898 to 0.3773 (+8.75% absolute / +30.2% relative, Acc@0.5 0% $\rightarrow$ 18.2%; up to +50.4% IoU on narrow stemware in 008.png and 0.9866 vs 0.7997 simulated), (2) Per-edge normalized entropy ($\tilde{H}_m = H_m / \ln 21$) spikes 1.37× on occluded box edges (0.6810 vs 0.4970 live; 2.86× simulated), and (3) DETR-style parallel object query slots prevent duplicate collapse via bidirectional self-attention. templates/multimodal/bbox_localization.json.tmpl
templates/multimodal/bbox_multi_object_detr.json.tmpl
templates/multimodal/bbox_multi_object_set.json.tmpl
benchmarks/bbox_suite.jsonl
dgem bench-bbox --annotate
results_bbox_cloudrun.json
results_bbox_simulated.json
✅ Completed
EXP-10 Listwise Diffusion Canvas Reranking, Softmax Expectation & RAG Poison Quarantine Evaluate 10 candidate passages (doc_01..doc_10) + 2 RAG security/abstention gates (answer_present, poisoned_passage) simultaneously in 1 forward pass (12 slots, ~138 ms effective/passage on Cloud Run 1× L4). Continuous Softmax Expectation ($\hat{r}i = \sum{g=0}^3 g \cdot p_{i,g}$) cuts Exact Tie Rate from 70.0% (discrete argmax) to 0.0%, lifting nDCG@10 from 0.8416 to 0.9265 (+8.49 pts) and MRR@10 from 0.7407 to 0.9444 (+20.37 pts), with 100.0% NevIR negation accuracy, +0.7533 FollowIR p-MRR policy steerability, and 100.0% prompt-injection quarantine. templates/rerank/listwise_decision_rerank.json.tmpl
templates/rerank/pointwise_rerank.json.tmpl
benchmarks/rerank_suite.jsonl
dgem bench-rerank
results_rerank_cloudrun.json
✅ Completed (Analysis)
EXP-11 JevBench v1.3.1 & v1.4 4-Axis Parity, Harmonic Scoring, Slot Temperature Calibration & Upstream Sync Upgrade dgem with JevBench v1.3.1 & v1.4's 4-Axis Scorecards (Chance-Corrected Intelligence, 10-Bin ECE + Soft TVD Calibration, Speed with 2×+0.15s load factor adjustment, Cost, and v1.4 Harmonic Mean Composite with 3 gates $&lt;50.0$) and Post-Hoc Slot Temperature Scaling ($p_k(T) = p_k^{1/T} / \sum p_j^{1/T}$). On Cloud Run EXP-04, T* = 1.35 cuts 10-bin ECE by 56.2% (0.0745 $\rightarrow$ 0.0326), lifting Calibration from 82.67 to 88.18 (76.79 Composite). On EXP-05, the Entropy-Gated Cascade scores 70.16 Composite (98.0% raw / 97.17% chance-corrected) vs. standalone gemini-3.8-flash at 62.76 (56% lower cost). Syncs and replays all 231 JevBench public tasks (75.70 v1.3.1 Composite $\rightarrow$ Rank #1, surpassing #1 Hopper [75.40]; 71.95* v1.4 Harmonic Estimate). Supports --compare-leaderboard against official v1.4.2 rows (Plumb-4B, decider-4b v2, djev, OpenJev). templates/jevbench_generic.json.tmpl
benchmarks/jevbench/jevbench_public.jsonl
benchmarks/jevbench/manifest.lock.json
benchmarks/jevbench/leaderboard_v14.json
dgem bench-jev --sync --check-upstream
dgem bench-jev --compare-leaderboard
results_djev_upstream_calibrated.json
✅ Completed (Analysis)
EXP-12 jev-decision-index (apolinario/decision-index) 5-Area Benchmark, Wide-Canvas Adapter & /v1/systemone Protocol Evaluate dgemma across the 5-Area Decision Index (19 scored panel benchmarks + 3 wide-option benchmarks). Eliminates HTTP 422 capacity rejections (0.0 score penalty on M > 10 slots in ContractNLI/BRIGHT/ToolRet and K > 26 options in API-Bank/BANKING77/CLINC150) via Multi-Slot Canvas Batching + 2-Stage Bracket Tournament Routing (T*=1.25), lifting structural coverage from 72.7% to 100.0% and Headline Decision Index from 76.67 to 98.89 (+22.22 pts, Skill 98.33, ECE 0.0371). Projects to Rank #3 / 49 Overall (#1 Diffusion, 46.90 balanced_skill) as Pure System-1 (0% LLM) and Rank #1 / 49 Overall (53.94 balanced_skill, beating TypeSafe Jev 1.13.0 [51.67]) with EXP-05/11 Entropy-Gated Cascades (#1 / 31 on v0.1 at 56.51 raw). benchmarks/decision_index/panel_suite.jsonl
pkg/decisionindex/engine.go
dgem bench-decision-index --compare-naive
dgem bench-decision-index --serve-systemone :8095
results_decision_index_cloudrun.json
✅ Completed (Analysis)
EXP-13 Permutation Sensitivity, Content-Free Null-Prior De-Biasing & $O(1)$ Dual-Mirror Canvas Calibration Quantifies how option order distorts single-pass confidence. Content-free probes show a strong Slot-'A' preference ($p_0 = [88.3%, 11.7%]$ for $K=2$, $[78.3%, 8.6%, 13.1%]$ for $K=3$, $[49.3%, 4.8%, 16.3%, 29.5%]$ for $K=4$). On 16 synthetic items (baseline 16/16 correct): (1) Null-Prior De-Biasing (13B) cuts Brier 0.0173 $\rightarrow$ 0.0017 but raises cyclic flip rate 12.5% $\rightarrow$ 25%; (2) Cyclic JSD (13A) is ~156× higher on the 4 ambiguous items, where reordering flips 2/4 answers; (3) Dual-Mirror (13C) reads forward + reversed slots in one pass (124.7 vs 128.8 ms) and flags perm_08 (99.9%, $\tilde{H}=0.007$, Mirror TVD = 0.258) but misses perm_06 (TVD 0.0006) and worsens Brier (0.0410). See the plain-English summary in Confidence Beyond Shannon (IDC). benchmarks/permutation_suite.jsonl
pkg/permutation/permutation.go
dgem bench-permutation
results_permutation_cloudrun.json
✅ Completed (Analysis)
EXP-14 Same-Session IDC Re-run on Vertex AI (versioned runs) Pre-registered follow-ups PROP-01/02/05/10 plus an EXP-13 reproduction, run in one session with versioned receipts. Findings: ±1–2 item noise floor; null-prior helps on the 50-item suite (Brier 0.147 vs 0.175–0.193) but not on JevBench (186 vs 187, Brier worse); the dual-mirror slot id __mirror_rev degraded readings (JevBench 187 → 145; fixed to __rev), and a same-canvas mirror still lowers the forward reading; held-out temperature scaling cuts ECE 24–33% on 231 JevBench items but not on 50 items; entropy cascade reaches 221/231 at 39% escalation. Details benchmarks/calibration_suite.jsonl, benchmarks/jevbench/jevbench_public.jsonl, benchmarks/permutation_suite.jsonl scripts/run_idc_rerun.sh
scripts/analyze_idc.py
benchmarks/runs/20260925-vertex-idc*
✅ Completed
EXP-15 Letter Collision in the Dual-Mirror (PROP-11) Same-canvas mirror damage comes from shared letters: on JevBench (231, Vertex G4, 3 baselines 182–189), an identical copy (184) and a digit-labelled reversed slot (182) are within noise, while the letter-labelled reversed slot drops to 154; forward answers change on 67% of same-letter items vs 6% otherwise (p≈1e-19). Details benchmarks/jevbench/jevbench_public.jsonl dgem bench-jev --dual-mirror --mirror-mode
scripts/analyze_idc.py collision
benchmarks/runs/20260925-prop11-letter-collision
✅ Completed
EXP-16 Slot Names Are Part of the Prompt (PROP-16) Single-slot ids (q1, random, mirror, check) are within the baseline band (183–189) on JevBench; a loaded second-slot id (__mirror_rev, identical options so no letter collision) drops accuracy to 162 (second slot 135) vs 181 for __rev. Style rule: neutral ids in multi-question schemas. Details benchmarks/jevbench/jevbench_public.jsonl dgem bench-jev --slot-id
benchmarks/runs/20260926-prop16-slot-names
✅ Completed
EXP-17 Separate-Pass Mirror (PROP-12) Forward vs. separately-run reversed pass on JevBench (231, Vertex G4): order disagreement relates to errors beyond hesitation (mean partial Spearman 0.133, CI [0.025, 0.240]), but two forward passes show a similar non-significant value (0.083) and CV AUROC improves by ≤ 0.016 at double the cost. Details benchmarks/jevbench/jevbench_public.jsonl dgem bench-jev --flip-options
scripts/analyze_idc.py separate-pass
benchmarks/runs/20260926-prop12-separate-pass
✅ Completed
EXP-18 Judge Capabilities With and Without an Autoregressive Autorater (mizan Experiment 07) Gold-scored comparison over 13 human-labelled judge suites (n=100 each) and 650 computation items, rerun on serving image v0.1.0 (2026-09-29) and replicated from 2026-09-26: v0.1.0 shows no regression (67.0 → 67.9%, p=0.23). DiffusionGemma matches gemini-3.8-flash on safety (80 vs 76), faithfulness (81 vs 78) and RewardBench with a mirror read (90 vs 90); trails on toxicity (81 vs 87.9, p=0.065) and LLMBar instruction following (79 vs 94, p=0.0015); MT-Bench expert preference favours Gemini by 5–10 points on both dates. A fixed 35%-hesitation cascade reaches 81/86 on safety/faithfulness with 8–32% escalation; p50 105 ms (G4) vs 3,125 ms. No-model metrics: mizan local = Vertex on exact/tool/trajectory (350/350). Speed/cost: one G4 replica ~83 items/s (0 errors sustained); ~$0.02 per 1k judgements fully utilized vs $1.06–2.25 (3.8-flash) and ≥$10 (Vertex predefined metrics); live cascade matches the simulation. Details mizan judge-eval pack; suites built by mizan scripts/judge-eval/build_suites.py (BeaverTails, ToxicChat, HaluBench, MT-Bench human, RewardBench, HelpSteer2, SummEval) mizan eval compare-engines
mizan scripts/judge-eval/run_all.sh
✅ Completed

3. Synthesis of Completed Findings (EXP-01 – EXP-04)

3.1 Why DiffusionGemma is a "Decision Model" (EXP-01 & EXP-04)

Across 280+ evaluated cases (EXP-01 through EXP-04), DiffusionGemma demonstrates three properties that distinguish a Decision Model from both autoregressive generative LLMs and discriminative encoder classifiers:

  1. Zero-Shot Policy-as-Template (0s Policy Iteration): Adding a new governance domain (AgentDrift behavioral drift, LLM-AggreFact RAG grounding, deepset/prompt-injections) required zero gradient updates, zero labeled training splits, and zero output regex parsers. Writing a .json.tmpl file immediately turned the 9B model into a 100%-accurate classifier across those four benchmarks (15/15 combined in EXP-04).
  2. Joint Multi-Slot Co-Adaptation in $O(1)$ Forward Pass: In EXP-01 (support_triage, code_review, security_audit) and EXP-04 (AgentDrift), a single forward pass (458.9 ms on 1× L4) evaluates 3 to 5 orthogonal decision dimensions simultaneously (drift_detected + drift_severity + remediation_action). Because the [MASK] canvas is bidirectional, the severity and remediation slots attend directly to the detection slot in the same pass.
  3. Intrinsic Epistemic Calibration (EXP-04 ChaosNLI 8.0× Entropy Multiplier): Traditional neural classifiers suffer from overconfidence on out-of-distribution or genuinely ambiguous inputs. In EXP-04, evaluating ChaosNLI (where 100 human annotators rated each premise/hypothesis pair) proved that DiffusionGemma's restricted-softmax Shannon entropy $H$ scales monotonically with human disagreement:
    • High human consensus (low-entropy): 100.0% accuracy, $H = 0.0744\text{ nats}$ (1.0× baseline)
    • Moderate split (ambiguous): $H = 0.4061\text{ nats}$ (5.5× entropy increase)
    • Near-uniform 3-way human split (high-entropy): $H = 0.5932\text{ nats}$ (8.0× entropy increase)
    • Adversarial multi-hop reasoning (ANLI-R3): When single-pass think: 0 readout fails (0/3), slot entropy automatically spikes to $H = 0.4104\text{ nats}$ (5.5×), providing an unmistakable mathematical trigger ($H &gt; 0.30\text{ nats}$) to escalate to a reasoning pass (EXP-05).