A regime-aware tabular fatigue benchmark for 22 MW floating offshore wind turbine towers.
FLOATBench is a public benchmark for surrogate modeling of FOWT
tower fatigue. It pairs 582,120 section-level fatigue damage
labels across three 22 MW floating-tower geometries with a
regime-aware evaluation protocol that stratifies test points into
in-train, interpolation, and extrapolation regions of the joint
wind/wave operating envelope. The dataset is hosted on Hugging Face
at DeCoDELab/FLOATBench;
this repository contains the benchmark code, evaluation harness, and
scripts to reproduce the paper results.
- João Alves Ribeiro (corresponding, MIT) — jpar@mit.edu
- Bruno Alves Ribeiro (TU Delft & Brown University)
- Francisco Pimenta (University of Porto)
- Sérgio M. O. Tavares (University of Aveiro)
- Faez Ahmed (MIT) — faez@mit.edu
FLOATBench is presented in the following paper, which fully describes the dataset, the regime-aware partition, and the evaluation protocol: FLOATBench: A Dataset and Benchmark for Floating Offshore Wind Turbine Tower Fatigue (under review at the NeurIPS 2026 Datasets and Benchmarks Track).
Across up to 96 tabular surrogates per tower (E1/E2) and up to 63 per fold (E3) — 735 trained surrogates in total — the regime-aware protocol reveals rank shifts between global and extrapolation performance that random-split leaderboards systematically miss, and a related rank inversion appears under cross-tower transfer. The release of the dataset, evaluation harness, and trained surrogates establishes common ground for adjudicating competing tabular surrogates on this domain.
-
Dataset. 582,120 rows of section-level fatigue damage across
three 22 MW FOWT towers (
ref,opt1,opt2), derived from high-fidelity OpenFAST simulations over a$22 \times 7 \times 7$ wind/wave operating envelope with 6 turbulence seeds and 30 tower sections per geometry. -
Regime-aware split. Alpha-shape partition of the joint
wind/wave envelope that labels each test row as
In-train,Interpolate, orExtrapolateon both axes, populating the full 9-cell regime grid. - Benchmark protocols. Three levels: random validation (E1), within-tower regime-aware evaluation (E2), and cross-tower transfer (E3).
-
Reproducible harness. End-to-end CLI scripts for training
(AutoGluon), evaluation (per-section / per-regime metrics),
bootstrap leaderboards, cross-preset benchmark plots, and the
alpha-shape splitter — all driven by
--flagfileconfigs.
Metrics. Throughout, DEL is the Damage Equivalent Load and Rel L² the relative L² error. The headline metric is Rel L² on DEL, reported globally and per regime / per section.
floatbench/ Python package (training, evaluation, plots, splitter)
scripts/ Pipeline entry points — see "Scripts" below
docs/ Figures and assets used in this README
environment.yml Conda environment (Python 3.12 + GPU PyTorch)
requirements.txt Pinned runtime dependencies
Each folder under scripts/ is a self-contained pipeline
stage with its own run.py and a --flagfile config.cfg. Together
they cover the full FLOATBench workflow, from recovering the split to
the cross-preset benchmark figures:
scripts/split/— reproduce or customize the regime-aware train/test split from the grid IDs; writes the per-towertrain_damage.csv/test_damage.csv, diagnostic plots, andsplit_metadata.json.scripts/train/— train an AutoGluon predictor for one preset (best/extreme) on a tower's train CSV.scripts/test/— predict on the test set and score it with per-section and per-regime (In-train / Interpolate / Extrapolate) metrics.scripts/leaderboard/— build the bootstrap CI leaderboard tables over DEL (paper Table 2).scripts/benchmark/— merge presets into the cross-preset benchmark outputs (regime heatmaps, bump chart, family bars,model_pooltable).scripts/run_benchmark.py— one-shot orchestrator that chains training, evaluation, leaderboard, and benchmark for the within-tower (E2) and cross-tower (E3) experiments.
Recommended (conda, GPU):
git clone https://github.com/Joao97ribeiro/FLOATBench
cd FLOATBench
conda env create -f environment.yml
conda activate floatbenchThis installs Python 3.12, PyTorch 2.6+ with CUDA 12.4, and all
AutoGluon backends (LightGBM, CatBoost, XGBoost, FastAI, TabM,
TabPFN, Mitra) plus the splitter / plot helpers from
requirements.txt.
Alternative (pip, CPU or existing venv):
git clone https://github.com/Joao97ribeiro/FLOATBench
cd FLOATBench
pip install torch # any torch>=2.6,<2.10
pip install -r requirements.txtThe released CSVs, schema, and per-tower layout are documented in
the dataset README on Hugging Face:
DeCoDELab/FLOATBench.
The three towers are:
-
ref— IEA-22-MW reference tower (baseline) -
opt1— first redesign iterate (relaxed damage budget,$D \le 1.0$ ) -
opt2— final iterate ($D \approx 0.9$ , targeting$D \le 0.9$ )
The opt1 and opt2 geometries were produced by
FLOAT, the fatigue-aware
tower design-optimization framework that the ref tower is redesigned with.
# Option A: download with the HF CLI (one-time)
hf download DeCoDELab/FLOATBench --repo-type=dataset --local-dir=data
# Option B: load on-the-fly from Python
python -c "from datasets import load_dataset; \
ds = load_dataset('DeCoDELab/FLOATBench', 'ref'); print(ds)"After this you should have data/{ref,opt1,opt2}/{train_damage.csv, test_damage.csv, data.csv, metadata.json}. See the
dataset README
for the full schema and the regime-aware split definition.
# Smoke test (~10 min total: 2 min per train, 2 trains, leaderboard, benchmark)
python scripts/run_benchmark.py --experiment=within --tower=ref \
--time_limit=120
# Full reproduction of the paper, all 6 experiments (E2 + E3, ~48 GPU-hours)
python scripts/run_benchmark.py --experiment=all
# Custom training budget (e.g. 8 h per training instead of the 4 h default)
python scripts/run_benchmark.py --experiment=all --time_limit=28800--time_limit controls the AutoGluon training budget (in seconds) per
preset and per experiment. Default is 14400 (4 h, paper setting). Use a
small value (e.g. 120) for a quick smoke test, or a larger value to
push beyond the paper budget. Outputs land in
outputs/within/{ref,opt1,opt2}/ and outputs/cross/{ref,opt1,opt2}/,
each containing trained models, leaderboards with bootstrap CIs, and a
cross-preset benchmark folder.
The benchmarks were run on a single workstation; nothing in the
pipeline assumes a cluster. The defaults in scripts/train/config.cfg
expose every knob:
| Resource | Paper setting | Notes |
|---|---|---|
| GPU | 1 × NVIDIA (24 GB used; 16 GB is enough) | Used for AutoGluon NN tabular families and TabPFN. Tunable via --num_gpus / --num_gpus_per_fold. |
| CPU | 24 cores total, 12 per bagging fold | Tunable via --num_cpus / --num_cpus_per_fold. |
| RAM | ≈ 32 GB | Peaks during AutoGluon ensembling. |
| Disk | ~225 MB dataset + ~5–10 GB trained models | One AutoGluon predictor per preset/tower/experiment. |
| Wall-clock | 4 h per preset (paper budget) | Set by --time_limit; full reproduction (3 towers × 2 presets × {within, cross}) ≈ 48 h. |
CPU-only training works for --presets=best (tree ensembles only) but
is significantly slower for --presets=extreme and effectively
disables zeroshot_2025_tabfm (TabPFN).
After a full run the experiment root is laid out like this:
outputs/within/ref/
├── best/model/ AutoGluon predictor (best preset)
│ ├── autogluon_meta.json config + features used at fit time
│ ├── leaderboard.csv AG built-in leaderboard (val score)
│ ├── leaderboard_test.csv same leaderboard, scored on test set
│ ├── leaderboard_test_summaries/
│ │ ├── leaderboard_test_metrics.csv r2 / Rel L² damage + DEL
│ │ ├── leaderboard_test_groups.csv per-regime metrics (IT/IP/EX × wind/wave)
│ │ ├── leaderboard_test_sections.csv per-section metrics (1 row per model × section)
│ │ └── del/ bootstrap CI95 over DEL (paper Table 2)
│ │ ├── leaderboard_test_summary.csv point estimates
│ │ ├── leaderboard_test_summary_ci95.csv 95% bootstrap CIs
│ │ ├── leaderboard_test_percentiles.csv bootstrap percentiles
│ │ ├── leaderboard_test_regime_rel_l2.csv Rel L² DEL per regime
│ │ └── leaderboard_test_section_rel_l2.csv Rel L² DEL per section
│ └── models/<MODEL_NAME>/test/predictions.csv per-model raw predictions
├── extreme/model/ (same layout, extreme preset)
└── benchmark/ cross-preset merge (the headline outputs)
├── model_pool.csv paper Table 9 (rows = preset, cols = family)
├── leaderboard/
│ ├── ranking/
│ │ ├── bump_chart.png paper Fig. 6 (rank movement)
│ │ ├── scatter_global_vs_ex_ex_*.png global vs EX_EX cross-over
│ │ ├── scatter_sections_top_models_*.png per-section scatter (sec1 / sec30 / EX_EX)
│ │ └── predictions_report.log which models had predictions, which were auto-generated
│ ├── regimes/
│ │ ├── heatmap_groups_mre_del.png 3×3 regime heatmap (paper Fig. 5)
│ │ └── heatmap_9groups_mre_del.png 9-cell expanded heatmap
│ ├── extrapolation/
│ │ ├── bar_family_regime_mre_del.png per-family Rel L² across regimes
│ │ └── scatter_global_vs_ex_ex_*.png
│ └── comparison/
│ └── family_distribution_rel_l2_del.png distribution of Rel L² across families
└── leaderboard_test_summaries/ merged across both presets (same files as per-preset)
Most CSVs are flat tables ready for downstream analysis; columns are
self-describing (r2_damage, rel_l2_del, rel_l2_del_EX_EX,
rel_l2_del_section_<i>, …). The bump_chart.png,
scatter_global_vs_ex_ex_*.png and heatmap_groups_*.png reproduce
the paper's headline E2 / E3 figures.
# Train one preset
python scripts/train/run.py --flagfile=scripts/train/config.cfg \
--train_csv=data/ref/train_damage.csv \
--test_csv=data/ref/test_damage.csv \
--output_dir=outputs/within/ref/best
# Evaluate
python scripts/test/run.py --flagfile=scripts/test/config.cfg
# Bootstrap leaderboard (DEL only)
python scripts/leaderboard/run.py --flagfile=scripts/leaderboard/config.cfg
# Cross-preset benchmark (heatmaps, bump charts, model_pool table)
python scripts/benchmark/run.py --flagfile=scripts/benchmark/config.cfgThe release ships pre-split CSVs that match the paper Table F.1
training set. The same splitter, however, lets you build alternative
training envelopes for ablations: change which wind setpoints, wave
pairs or seeds are used for training by picking different grid IDs
on the
The defaults in scripts/split/config.cfg reproduce the paper split
byte-for-byte. Run it as is:
python scripts/split/run.py --flagfile=scripts/split/config.cfgTo explore an alternative envelope, copy scripts/split/config.cfg
and edit any of the --train_ws_ids, --train_hs_ids,
--train_tp_ids lines. Excerpt of the config:
# wind_speed_id values that go to train (excluded: 1, 8, 15, 22)
--train_ws_ids=2,3,4,5,6,7,9,10,11,12,13,14,16,17,18,19,20,21
# wave_hs_id values that go to train (excluded: 1, 4, 7)
--train_hs_ids=2,3,5,6
# wave_tp_id values that go to train (excluded: 1, 4, 7)
--train_tp_ids=2,3,5,6Or override on the command line (e.g., denser wave envelope):
python scripts/split/run.py --flagfile=scripts/split/config.cfg \
--train_hs_ids=1,2,3,4,5,6,7 \
--output_dir=outputs/split_denserEach run writes the per-tower train_damage.csv/test_damage.csv,
the train + test diagnostic plots, and a top-level
split_metadata.json with the grid summary and train-spacing
statistics.
The boundary between Interpolate and Extrapolate is fixed in the
released splitter (alpha-shape on the train hull plus a normalized
distance threshold of 0.5). To explore a different threshold,
instantiate floatbench.split.domain_groups.WindWaveDomainGrouper
directly with custom interp_edges / extrap_edges.
This lets you construct your own train/test splits without re-simulating any OpenFAST cases.
Within-tower (E2): the global rank-1 fails at the boundary. On every
tower, the AutoGluon default ensemble (WeightedEnsemble_L2) ranks first
globally yet is overtaken at the worst-case wind-and-wave extrapolation
cell (EX_EX) by a neural-network family the greedy selector systematically
excludes:
Cross-tower (E3): transfer is asymmetric. Training on a set that
includes the baseline ref generalises well to the perturbed geometries
(rank-1 Rel L² DEL of 0.067 / 0.098). Training without ref, however,
collapses to 0.423 — a 4–6× degradation:
Code released under the MIT License. Dataset on Hugging Face is released under CC-BY-4.0.
If you use FLOATBench in your work, please cite:
FLOATBench: A Dataset and Benchmark for Floating Offshore Wind Turbine Tower Fatigue. João Alves Ribeiro, Bruno Alves Ribeiro, Francisco Pimenta, Sérgio M. O. Tavares, Faez Ahmed. arXiv:2605.25717, 2026. https://arxiv.org/abs/2605.25717
BibTeX
@misc{ribeiro2026floatbenchdatasetbenchmarkfloating,
title={FLOATBench: A Dataset and Benchmark for Floating Offshore Wind Turbine Tower Fatigue},
author={João Alves Ribeiro and Bruno Alves Ribeiro and Francisco Pimenta and Sérgio M. O. Tavares and Faez Ahmed},
year={2026},
eprint={2605.25717},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.25717},
}For issues, questions, or feature requests related to FLOATBench: FLOATBench Issues.
The high-fidelity simulations underlying FLOATBench were produced with OpenFAST on the IEA-22-280-RWT reference floating wind turbine. The tabular surrogate pipeline relies on AutoGluon. We thank these communities for keeping the underlying tools open.





