A reliability-first machine-learning framework for heterogeneous biodiversity data integration
Research question: When can heterogeneous biodiversity data integration actually be trusted?
BioTrust-Fusion integrates structured North American Breeding Bird Survey (BBS) monitoring, opportunistic eBird observations, conventional environmental covariates derived from NASA AppEEARS, NLCD, and USGS 3DEP, and Earth-observation representations from Prithvi, TerraMind, and Clay.
The project is deliberately not a simple “which model wins?” benchmark. It asks whether apparent gains survive independent checks of source transport, geographic generalization, temporal generalization, uncertainty, support, sensitivity stability, response compatibility, and failure behavior.
The central empirical finding is:
Heterogeneous biodiversity data and Earth-observation foundation-model representations can provide genuine information gain, but the strongest method depends on which reliability dimension is examined.
BioTrust-Fusion therefore separates information gain from reliable inference.
| Finding | Evidence |
|---|---|
| Real ecological signal after integration | Red-eyed Vireo Anchored abundance: +1.171%/yr, 95% bootstrap CI [+0.204, +2.151], sign stability 0.996 |
| Primary uncertainty analysis | 500 / 500 clustered bootstrap replicates completed |
| Observer robustness | 23 / 24 endpoint directions preserved under observer-disjoint evaluation |
| Weight robustness | 11 / 12 Anchored endpoint directions preserved under transport-weight cap 20 |
| Most balanced EO signal | Prithvi: 50.9% geographic and 51.4% temporal non-tie win rates |
| Strongest spatial EO consistency | Clay: 58.2% geographic non-tie win rate |
| Important negative control / counterexample | TerraMind: comparatively benign temporal transport did not translate into ecological predictive gain |
| Core scientific result | Prediction, precision, and transport quality can disagree; none alone is sufficient to establish reliable integration |
Biodiversity inference increasingly combines data collected under very different sampling mechanisms.
BBS is structured monitoring: routes, protocols, repeated observations, and comparatively controlled sampling.
eBird is participatory monitoring: vastly larger coverage, but observation intensity, location choice, observer behavior, and sampling effort are not generated by the same process.
Earth-observation models introduce another layer of heterogeneity. Rich learned representations may improve prediction while simultaneously changing source overlap, leverage, or temporal stability.
BioTrust-Fusion evaluates these pieces within one common framework rather than equating better predictive performance with trustworthy ecological inference.
The real-data analysis covers:
- New York
- Ohio
- Pennsylvania
- Vermont
during the breeding-season window:
May 15 – July 10
Spatial analysis uses an EPSG:5070 25-km grid.
The broader study-region cell universe contains 241 analysis cells. Cross-source support filtering produces a smaller shared support set, and the frozen conventional ecological reporting target uses 52 cells × 43 dates/year.
BBS
- 719 acceptable route-years
- 68 routes
- 2010–2025 study period
- 2020 treated as structurally unsampled
eBird
- Cornell eBird EBD release
- NY / OH / PA / VT
- 547,183 eligible checklist events after the frozen filtering pipeline
After harmonization with the focal-species panel, the model adapter contained 2,739,510 event × species response rows.
BBS 2020 is never fabricated or used as an ecological evaluation year.
The project has two connected but scientifically distinct branches:
- Primary BBS–eBird integration
- Earth-observation representation experiment
Earth-observation embeddings are therefore not prerequisites for the primary BBS–eBird estimator comparison. They are a later representation experiment evaluated through the same reliability philosophy.
This figure compares Conventional environmental covariates, Prithvi, TerraMind, and Clay under the same reliability framework.
The EO experiment reveals distinct representation profiles:
- Prithvi provides the most balanced geographic and temporal predictive evidence.
- Clay provides the strongest geographic predictive consistency.
- TerraMind demonstrates that comparatively favorable transport behavior does not necessarily translate into ecological predictive utility.
- Conventional covariates remain the most defensible geographic transport reference.
The central EO finding is therefore not a universal encoder winner.
Instead, different environmental representations succeed or fail along different reliability dimensions.
BioTrust-Fusion preserves downloaded source material separately from all derived data.
Downloaded source archive
|
v
raw/
immutable source
|
v
extracted/
exact unpacked contents
|
v
filtering / QA
|
v
harmonization
|
v
analysis-ready data
BioTrust-Fusion uses two distinct environmental-data branches:
- Conventional environmental covariates used in the primary biodiversity integration framework and as the baseline environmental representation.
- Multispectral Earth-observation imagery used later for the Prithvi, TerraMind, and Clay representation experiment.
The conventional environmental representation was constructed from multiple climate, vegetation, land-cover, and terrain products.
NASA AppEEARS
-
Daymet V4 (
DAYMET.004), 2010–2025- precipitation
- solar radiation
- maximum temperature
- minimum temperature
- vapor pressure
-
MODIS Vegetation Indices (
MOD13Q1.061)- NDVI
- EVI
- vegetation-index quality
- pixel reliability
- composite day-of-year
Additional geospatial sources
- NLCD — land-cover composition
- USGS 3DEP — elevation and terrain information
These environmental products were spatially harmonized and aggregated to the frozen EPSG:5070 25-km analysis cells.
The resulting conventional environmental features were used for:
- source-propensity and transport modeling between BBS and eBird;
- ecological outcome modeling;
- geographic and temporal validation; and
- the baseline representation against which Earth-observation foundation-model representations were compared.
The foundation-model experiment used two multispectral satellite-imagery sources.
NASA Earthdata / LP DAAC — Harmonized Landsat Sentinel-2 (HLS v2.0)
- Collections: HLSL30 and HLSS30
- Surface-reflectance imagery
- QA source: Fmask
- Cloud, cloud-shadow, adjacent-cloud/shadow, and snow/ice pixels were excluded
- Prithvi input bands:
- B02
- B03
- B04
- B05
- B06
- B07
- Four breeding-season temporal slots were selected for each supported cell-year
- Selected imagery was cropped/resampled to 224 × 224 model-ready chips
The selected HLS imagery was passed through the frozen Prithvi-EO-2.0-300M-TL backbone to obtain 1024-D environmental representations.
Sentinel-2 L2A — Element84 Earth Search STAC
Sentinel-2 surface-reflectance scenes were independently discovered and quality-filtered using the Scene Classification Layer (SCL).
TerraMind used 12 Sentinel-2 bands and generated a 768-D representation.
Clay used the same matched Sentinel-2 scenes as TerraMind, with its sensor-specific 10-band input configuration, and generated a 1024-D representation.
Satellite imagery was not downloaded indiscriminately.
The EO pipeline first performed metadata-only scene discovery, followed by quality screening and deterministic scene ranking.
Candidate scenes were evaluated using:
- local QA-valid pixel fraction;
- cloud cover;
- temporal distance from the center of each breeding-season time bin; and
- deterministic item-ID tie breaking.
A minimum of 85% locally valid pixels was required.
Only the selected scenes were then downloaded and cached as reusable reflectance inputs for representation extraction.
The four environmental comparison arms were:
Conventional environmental features
vs
Conventional + Prithvi representation
vs
Conventional + TerraMind representation
vs
Conventional + Clay representation
This figure compares the Structured-only and Anchored-transport estimators across the primary ecological endpoints.
It combines:
- ecological trend estimates,
- 95% clustered-bootstrap uncertainty,
- Structured-versus-Anchored differences,
- transport diagnostics,
- direction agreement,
- sensitivity evidence.
The strongest Anchored ecological signal occurs for Red-eyed Vireo abundance:
+1.171%/yr, 95% bootstrap CI [+0.204, +2.151], with sign stability 0.996.
The broader result is that integrating eBird can change the ecological estimate itself rather than merely narrow uncertainty.
Four species define the primary ecological analysis:
- American Robin — Turdus migratorius
- Red-eyed Vireo — Vireo olivaceus
- Barn Swallow — Hirundo rustica
- Red-winged Blackbird — Agelaius phoeniceus
A fifth species is used as a prespecified low-information sensitivity benchmark:
- Red-shouldered Hawk — Buteo lineatus
Image provenance and licenses are documented in:
phase4_real_data/assets/species/SPECIES_IMAGE_CREDITS.md
The map shows the four-state study region, the 25-km analysis-cell footprint, and the eight geographic outer-validation blocks used to test spatial generalization.
The map is not decorative: it visualizes the spatial support on which the reliability analysis is built.
The publication-ready tables below summarize the main empirical evidence from BioTrust-Fusion.
This table compares Structured-only and Anchored-transport ecological estimates across the four primary species and three ecological components, including:
- annual trend estimates,
- 95% clustered-bootstrap confidence intervals,
- Anchored-minus-Structured differences,
- Anchored sign stability,
- direction agreement.
Machine-readable CSV:
Table1_Primary_BBS_eBird_Endpoint_Results.csv
This table compares Conventional, Prithvi, TerraMind, and Clay across geographic and temporal evaluation, including predictive consistency and source-transport diagnostics.
It makes the central EO result visible:
- Prithvi provides the most balanced geographic/temporal predictive evidence.
- Clay provides the strongest geographic predictive consistency.
- TerraMind demonstrates that favorable transport behavior does not automatically imply ecological predictive utility.
Machine-readable CSV:
Table2_EO_Representation_Benchmark.csv
This is the project-level synthesis table.
It brings together:
- BBS + eBird integration,
- Conventional environmental representation,
- Prithvi,
- Clay,
- TerraMind,
- the overall BioTrust-Fusion scientific contribution.
The table separates each arm's strongest empirical finding, robustness/generalization evidence, reliability boundary, and scientific contribution.
Machine-readable CSV:
Table3_BioTrust_Fusion_Master_Reliability.csv
Earlier compact machine-readable summaries are retained for reproducibility:
These supporting files are retained as audit artifacts; Tables 1–3 above are the publication-facing result tables.
Primary ecological uncertainty was evaluated using 500 clustered bootstrap replicates:
Attempted: 500
Completed: 500
Failed: 0
Interval: 95% percentile bootstrap CI
Direction was preserved for 23 of 24 endpoint comparisons.
The only sign change occurred for the Barn Swallow Anchored multiplicative component.
Capping transport weights at 20 preserved 11 of 12 Anchored endpoint directions.
Again, the only direction change occurred for the Barn Swallow Anchored multiplicative component.
The agreement between these two sensitivity experiments identifies a specific fragile result rather than hiding instability inside an overall average.
Attempted: 500
Completed: 410
Failed: 90
Failure rate: 18%
Because failure was non-negligible, survivor-only uncertainty intervals were withheld.
BioTrust-Fusion therefore treats computational failure itself as reliability evidence.
The Red-shouldered Hawk sensitivity benchmark completed 20 / 20 planned runs.
Ecological replicate values were not stored under that benchmark contract, so no unsupported confidence interval, standard-error, or sign-stability claim is reconstructed afterward.
The intended Earth-observation study period was 2017–2025.
Under the frozen support rule, the realized primary common-support period is effectively:
2018–2025
No spatial cell remains supported across every ecological reporting year required for a fixed-spatial-support 2018–2025 EO trend.
BioTrust-Fusion therefore does not construct a longitudinal EO trend by silently changing the target spatial population across years.
Instead, the EO experiment asks:
Does richer environmental representation provide useful and reliable information on the exact common support where comparison is scientifically valid?
BioTrust-Fusion treats reliability as a combination of evidence:
Ecological / predictive information
+
Source transport
+
Geographic generalization
+
Temporal generalization
+
Response compatibility
+
Uncertainty
+
Common support
+
Sensitivity stability
+
Failure behavior
|
v
RELIABILITY EVIDENCE
The project provides several concrete empirical lessons:
-
More data can change the ecological conclusion. Anchored eBird integration can change trend magnitude or direction rather than merely narrow uncertainty.
-
Precision is not equivalent to absence of bias. A more precise estimator can still rely on unstable source transport.
-
Richer EO representations are not uniformly better. Prithvi, TerraMind, and Clay exhibit different predictive and transport profiles.
-
Geographic success does not guarantee temporal success. Clay provides the clearest example.
-
Transport adequacy does not guarantee ecological usefulness. TerraMind provides the complementary counterexample.
-
Sensitivity analysis can identify exactly where inference is fragile. Barn Swallow’s Anchored multiplicative endpoint is the only direction-changing result under both observer-disjoint and capped-weight checks.
-
Computational failure is information. The route-cluster bootstrap shows why failed replicates should not simply be discarded.
-
Support defines what can legitimately be estimated. Available embeddings alone do not justify a longitudinal ecological trend when fixed spatial support does not exist.
BioTrust-Fusion represents approximately 300 person-hours of research and development, including biodiversity data engineering, eBird filtering, BBS harmonization, support construction, source-propensity modeling, geographic and temporal validation, bootstrap uncertainty, sensitivity testing, Earth-observation acquisition and QA, Prithvi/TerraMind/Clay inference, representation reduction, reproducibility engineering, QC, and final evidence synthesis.
This refers to development effort, not model runtime.
phase4_real_data/outputs/figures/
Figure1_BioTrust_Fusion_Master_Workflow.png
Figure2_EO_Final_Evidence.png
Figure3_Primary_Integration_Reliability.png
Figure4_Study_Species_Panel.png
Figure5_Study_Region_Map.png
Table1_Primary_BBS_eBird_Endpoint_Results.png
Table2_EO_Representation_Benchmark.png
Table3_BioTrust_Fusion_Master_Reliability.png
PDF versions are also available for publication-quality rendering.
phase4_real_data/outputs/modeling/
biotrust_final_evidence_synthesis.json
biotrust_final_species_reliability_summary.csv
phase4d34d_primary_endpoint_evidence.csv
Additional project documentation is available in:
docs/METHODS.md
docs/RESULTS.md
docs/RELIABILITY_FRAMEWORK.md
BioTrust-Fusion/
|
+-- README.md
+-- configs/
+-- docs/
+-- scripts/
+-- src/
+-- tests/
|
+-- phase4_real_data/
+-- assets/
+-- configs/
+-- scripts/
+-- src/
+-- tests/
\-- outputs/
+-- environment/
+-- figures/
+-- modeling/
+-- reports/
\-- tables/
Before the final real-data analysis, BioTrust-Fusion used controlled synthetic experiments to implement and stress-test cross-fitted Double Machine Learning (DML) trend estimators, including a transport-aware variant.
These early experiments informed the project's later emphasis on grouped cross-fitting, out-of-fold estimation, leakage control, and explicit source-transport diagnostics.
The final real-data analysis does not present its primary estimator as a fully orthogonal DML estimator. Primary inference is reported using Structured-only and Anchored-transport estimators, and no claim of Neyman orthogonality to source-propensity estimation is made.
BioTrust-Fusion was built around explicit experiment contracts rather than retrospective model selection.
Core principles include:
- immutable raw-source preservation
- frozen study definitions
- auditable filtering
- explicit support rules
- grouped cross-fitting
- fold-local preprocessing
- no global representation fitting
- common evaluation universes
- representation-specific propensity refitting
- prespecified sensitivity analyses
- reproducible QC artifacts
- failure-aware uncertainty
- explicit separation of predictive gain from reliability
The final computational evidence pipeline is frozen.
Large, restricted, or credential-dependent datasets are intentionally excluded from public version control.
These include:
- raw eBird source data
- restricted biodiversity source extracts
- downloaded HLS imagery
- downloaded Sentinel-2 imagery
- raster caches
- model checkpoints
- full embedding stores
- temporary processing products
- environment-specific credentials
The public repository instead preserves the code, frozen configurations, tests, audit artifacts, aggregate evidence, figures, tables, and documentation needed to understand the computational workflow.
BioTrust-Fusion found real ecological and predictive information gain.
It also demonstrated why those gains cannot be interpreted from prediction or precision alone.
- Red-eyed Vireo demonstrated a strong Anchored ecological signal.
- Prithvi provided the most balanced EO predictive evidence.
- Clay provided the strongest geographic predictive consistency.
- TerraMind demonstrated that apparently favorable transport does not necessarily imply ecological utility.
- Sensitivity analysis identified both broad stability and one specific fragile Barn Swallow endpoint.
Trustworthy biodiversity integration is conditional. BioTrust-Fusion provides a reproducible way to determine which apparent gains survive the checks required for scientifically defensible inference.







