Conversation
Persist successful run IDs and allow completed experiments to be tracked without retraining after a temporary MLflow outage. Co-authored-by: Cursor <cursoragent@cursor.com>
…tion-screen Add XC optimizer and regularization screen
Test shared vertical pressure-profile CNN experiment
…mask Test physics-aware AGL pressure-profile representation
…mask-confirmation Generalize paired-seed confirmation and close Conv1D profile experiments
Benchmark Jev 1.13 on the XC temporal holdout
Document the completed raw-weather and production-site Jev 1.13 ablation, preserve the reproducible benchmark implementation, and close the Jev experiment line.
Record the failed seed-42 screen in ADR 0017 and keep the implementation as a reproducible negative experiment.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Introduces a model-family-agnostic ML experimentation area under
ml/, with S2S as the first complete task, and migrates the production XC workflow behind explicit benchmark, reproducibility, ONNX-parity and promotion gates.The XC migration has since progressed beyond parity reconstruction into a reproducible architecture/optimization program on the fixed
xc-temporal-2024-jan-nov-v1benchmark.Included
glideator_ml.xcAccepted XC decisions so far
8192for RTX 3090 architecture experiments (ADR 0007)[64, 32]after paired seeds 42–46 (ADR 0011)The promoted smaller encoder beat the previous no-CrossNet control on all five seeds for macro BCE, macro Brier and macro ROC-AUC while reducing the model from roughly 64k to 48.5k trainable parameters. Monotonicity violations increased and remain a reported secondary diagnostic.
Current conventional XC baseline
[64, 32][64, 32]328192Canonical config:
ml/configs/xc/baselines/conventional_mlp.yaml.Active experiment
Site-embedding size screen at seed 42: dimensions
8,16,32(control), and64. Only the embedding dimension changes; tests guard the comparison contract. At most one challenger will be promoted to paired seeds 42–46.After site embedding size, the planned conventional optimization sequence is learning rate → dropout/AdamW weight decay → ReLU vs SiLU → freeze the tuned conventional MLP benchmark. Weather-structured architectures come after that.
Production boundary
Before XC cutover
net/path only after cutoverNot included
This PR remains intentionally draft while XC migration and model selection are still active.