Skip to content

About

A small DDPM built from scratch, generating a distribution of next-month temperature anomalies instead of one point estimate. Calibration is reasonable, close to an 80% coverage target. The central prediction overreacts to its input, losing to a linear baseline. A fix attempt made this worse, not better, reported directly.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

A Small Diffusion Model for Global Temperature Anomaly Uncertainty

A demonstration of denoising diffusion (DDPM) mechanics, built by hand, tested for honest calibration rather than point accuracy.

Overview

Every model built earlier in this project series outputs one predicted number, with uncertainty handled separately through a prediction interval or a permutation test. This notebook tries a different approach, training a small diffusion model to generate a distribution of plausible next month values, conditioned on the current month, instead of a single estimate with an interval added afterward.

The dataset is NASA GISTEMP's global monthly temperature anomaly record, 1880 to 2026, a different dataset from the Arctic sea ice work elsewhere in this series, used here because a diffusion model needs more data than the 45 or so annual points used in the earlier projects. The question asked is not primarily about temperature itself, it is whether this kind of model can produce a genuinely well calibrated distribution of outcomes, checked directly rather than assumed.

Data

1758 monthly anomaly values, 1880 to 2026, pulled directly from NASA GISTEMP.

Raw monthly temperature anomaly, 1880 to 2026 The full raw record. Flat and noisy from 1880 to roughly the 1970s, then a clear, accelerating rise from there to the present.

A 12-month centred rolling mean is subtracted to remove that trend, since a straight-line detrend would be misspecified given its accelerating, non-linear shape, the same kind of issue the Trend-Break project found elsewhere in this series. 1747 usable monthly residuals remain, standard deviation 0.097°C.

Residual after trend removal The same record with the rolling trend subtracted. No visible long-term drift left, just month to month variability, which is what the model is actually trained on.

How the model works

1746 conditioning pairs are built, this month's residual predicting next month's residual. A two-layer denoising network is built from scratch in PyTorch, with a sinusoidal time embedding for the diffusion step, following the standard DDPM forward and reverse process, 200 steps, linear beta schedule.

One value, increasingly noised through the forward process A single real value, shown at five points as it is progressively noised. By the final step it is indistinguishable from random noise, this is the fixed process the network is trained to reverse.

Trained for 300 epochs, batches of 128, one averaged loss per batch, avoiding the exact gradient-accumulation bug caught and fixed in the earlier GNN project.

Training loss over time Loss drops sharply in the first few epochs, then plateaus around 0.6. The plateau is expected, not a sign of failure, since the network is predicting randomly sampled noise, and even a perfect model cannot drive that loss to zero.

What the model actually learned

Before evaluating the model, the real relationship in the data was checked directly.

Relationship between consecutive month residuals This month's residual against next month's, across all 1746 pairs. A real but weak relationship, correlation 0.226, visible as a faint diagonal lean in an otherwise shapeless cloud.

Point prediction. The model's generated mean correlated with the conditioning value at 0.7895, well above that true 0.226. It responds to each month's value roughly 3.5 times more strongly than the data supports, and loses to a simple linear baseline on accuracy.

Generated distribution for one held-out test example 200 generated samples for a single test month. The spread looks reasonable, but the distribution's centre sits further from zero than the weak true relationship would justify.

Method MAE
Naive (predict 0) 0.8648
Linear baseline (uses the true 0.226 correlation directly) 0.8454
Diffusion model, mean prediction 0.8620

Calibration, checked separately from the point estimate, using an 80% coverage target, the same target used in the causal project's own conformal prediction interval.

80% interval coverage check 150 held-out test points, sorted by conditioning value. Blue band is each point's 80% generated interval; red dots are the actual values that fell outside it. Misses are scattered across the range rather than concentrated at one extreme, empirical coverage came out at 76.0%, close to target.

Calibration and the honesty of the central estimate turned out to be two separate properties of the same model. This one has more of the first than the second, a wide-enough interval to catch the true value most of the time, wrapped around a central guess that overreacts to its input.

Testing a fix, and being wrong about why

The working hypothesis was that 300 epochs let the network overfit the weak conditioning relationship. A second model was trained identically, for 80 epochs instead of 300, to test this directly.

Training loss, 300 vs 80 epochs Both runs reach the same loss plateau, confirming 80 epochs is enough to match 300 epochs' training loss, so any difference in the results below isn't just a matter of the shorter run being undertrained.

The result was the opposite of what the hypothesis predicted.

Metric 300 epochs 80 epochs True value
Correlation with conditioning value 0.7895 0.9157 0.2257
MAE, mean prediction 0.8620 0.8537 —
Coverage vs 80% target 76.0% 78.0% —

Fewer epochs made the conditioning overuse worse, not better, even though point accuracy and calibration both improved slightly. The original hypothesis does not hold up against this evidence. A plausible alternative, not confirmed here, is that the conditioning value is the easiest, most immediately available signal in the input, and a network early in training may lean on it as a shortcut before it has learned the harder, more diffuse structure of the actual noise, with longer training being what eventually lets it settle into a more measured response.

Summary: calibration is reasonable, the central estimate is not All three findings side by side. Left: both training runs overshoot the true correlation, the shorter run more so. Middle: neither diffusion configuration beats the simple linear baseline on point accuracy. Right: both configurations land close to the 80% calibration target.

Conclusion

This notebook trained a small conditional diffusion model on real global monthly temperature anomaly data to generate a distribution of plausible next-month values rather than a single point estimate. Calibration is reasonable, close to an 80% target across two different training configurations. The central prediction is not reliable, responding to the conditioning value three to four times more strongly than the real, weak correlation supports, and losing to a simple linear baseline on point accuracy in both configurations tried.

Adjusting the number of training epochs did not fix this, and moved the result in the opposite direction from what was expected, which is itself a more informative finding than either configuration's individual numbers. The honest conclusion is that this model demonstrates the diffusion mechanism working correctly and produces sensible, reasonably calibrated distributions, but has not been shown to extract a genuinely well-calibrated conditional mean at this scale, with this amount of data, in this amount of exploration.

Limitations

  • 80 years of the 146-year record predate reliable global instrumental coverage at the density used in later decades; no adjustment was made for this
  • Only two training durations were compared; the true cause of the conditioning overuse remains an open question, not a solved one
  • The conditioning signal used is only the immediately preceding month; no attempt was made to condition on a longer history
  • The most direct next step is testing whether an explicit regularisation term, rather than epoch count alone, controls the model's over-reliance on the conditioning input

Reproducing

pip install torch numpy pandas matplotlib

Data is pulled directly from NASA's public GISTEMP archive, no download or account required. No GPU needed or useful at this scale. Full training and evaluation runs in a few minutes on CPU.

About

A small DDPM built from scratch, generating a distribution of next-month temperature anomalies instead of one point estimate. Calibration is reasonable, close to an 80% coverage target. The central prediction overreacts to its input, losing to a linear baseline. A fix attempt made this worse, not better, reported directly.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages