Skip to content

Build ML experiment platform and start XC migration - #115

Draft
janhelcl wants to merge 276 commits into
mainfrom
feature/ml-experiment-platform
Draft

janhelcl wants to merge 276 commits into
mainfrom
feature/ml-experiment-platform

Conversation

@janhelcl

@janhelcl janhelcl commented Aug 30, 2026 •

Copy link
Copy Markdown
Owner

Summary

Introduces a model-family-agnostic ML experimentation area under ml/, with S2S as the first complete task, and migrates the production XC workflow behind explicit benchmark, reproducibility, ONNX-parity and promotion gates.

The XC migration has since progressed beyond parity reconstruction into a reproducible architecture/optimization program on the fixed xc-temporal-2024-jan-nov-v1 benchmark.

Included

  • deterministic S2S benchmark identity, walk-forward evaluation, contrastive/SVD experiments, scorer-aware artifacts, MLflow provenance, docs and CI coverage
  • production XC architecture, preprocessing and objective migrated under glideator_ml.xc
  • TorchRec-free legacy-compatible CrossNet and exact served-model weight/parity support
  • leak-free XC model selection: pre-2023 fit window, calendar-2023 validation, untouched 2024-01-01 through 2024-11-30 evaluation
  • benchmark/data/evaluation fingerprints and boundary guards
  • config-driven deterministic XC training, evaluation, artifact creation and MLflow tracking/backfill
  • ONNX export and PyTorch ↔ ONNX Runtime numerical parity checks
  • one-command served-reference → candidate → promotion workflow
  • reusable paired multi-seed confirmation runner with shared prepared-data snapshot and paired metric summaries
  • TabPFN-3 frontier tabular baseline in both ordinal-multiclass and independent-binary formulations
  • documented architecture decisions and reproducible experiment configs

Accepted XC decisions so far

  • batch size 8192 for RTX 3090 architecture experiments (ADR 0007)
  • remove CrossNet from the candidate baseline (ADR 0008)
  • retain independent multilabel output heads; hard-monotonic heads regress predictive quality (ADR 0009)
  • do not replace the main XC model with TabPFN-3 (ADR 0010)
  • promote the smaller shared per-time encoder [64, 32] after paired seeds 42–46 (ADR 0011)

The promoted smaller encoder beat the previous no-CrossNet control on all five seeds for macro BCE, macro Brier and macro ROC-AUC while reducing the model from roughly 64k to 48.5k trainable parameters. Monotonicity violations increased and remain a reported secondary diagnostic.

Current conventional XC baseline

  • no CrossNet
  • raw per-time bypass retained
  • shared per-time encoder [64, 32]
  • fusion [64, 32]
  • site embedding dimension 32
  • independent eleven-output sigmoid head
  • batch size 8192

Canonical config: ml/configs/xc/baselines/conventional_mlp.yaml.

Active experiment

Site-embedding size screen at seed 42: dimensions 8, 16, 32 (control), and 64. Only the embedding dimension changes; tests guard the comparison contract. At most one challenger will be promoted to paired seeds 42–46.

After site embedding size, the planned conventional optimization sequence is learning rate → dropout/AdamW weight decay → ReLU vs SiLU → freeze the tuned conventional MLP benchmark. Weather-structured architectures come after that.

Production boundary

  • backend serving remains unchanged and still uses the checked-in production ONNX artifact
  • no automatic model replacement or deployment
  • experiment promotion does not itself change serving

Before XC cutover

  • finish conventional optimization
  • compare the final candidate against the served reference under the strict promotion contract
  • replace the backend artifact only after an explicit promotion decision
  • retire the legacy notebook/net/ path only after cutover

Not included

  • no D2D migration yet
  • no model registry or automatic retraining/deployment

This PR remains intentionally draft while XC migration and model selection are still active.

janhelcl and others added 30 commits August 30, 2026 17:07
Persist successful run IDs and allow completed experiments to be tracked without retraining after a temporary MLflow outage.

Co-authored-by: Cursor <cursoragent@cursor.com>
…tion-screen

Add XC optimizer and regularization screen
Test shared vertical pressure-profile CNN experiment
…mask

Test physics-aware AGL pressure-profile representation
…mask-confirmation

Generalize paired-seed confirmation and close Conv1D profile experiments
Benchmark Jev 1.13 on the XC temporal holdout
Document the completed raw-weather and production-site Jev 1.13 ablation, preserve the reproducible benchmark implementation, and close the Jev experiment line.
Record the failed seed-42 screen in ADR 0017 and keep the implementation as a reproducible negative experiment.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant