You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This repository is a supplementary to the manuscript "Integrating Simulations and Observations: A Foundation Model for Estimating Aerosol Mixing State Index"
The objective of this project are:
Pre-train a foundation model for aerosol mixing state prediction using PartMC-MOSAIC simulation data.
Fine-tune the pre-trained foundation model with MEGAPOLI observational data.
Analyze the impact of data scarcity on the performance of the fine-tuned model and input feature importance.
MEGAPOLI data: MEGAPOLI observational data will be made available on request.
Fine_tuned_Results_different_data_szie
Folder
Comments
How to get it?
Fine_tuning_XX%Data.csv
Chi estimation results from fine-tuned foundation model, here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning training dataset)
Chi estimation results from fine-tuned foundation model ('Swapped Temporal Order' experiment), here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning testing dataset), model performance evaluated on fine-tuning training dataset
Chi estimation results from AutoML, here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning training dataset)
Chi estimation results from AutoML, here XX means training dataset is XX fraction of total MEGAPOLI data ('Swapped Temporal Order' experiment) (XX *2 fraction of fine-tuning testing dataset), model performance evaluated on fine-tuning training dataset
Chi estimation results from Linear regression, here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning training dataset)
Chi estimation results from Linear regression ('Swapped Temporal Order' experiment), here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning testing dataset), model performance evaluated on fine-tuning training dataset
Used 20% PartMC data to train the pre-trained foundation model, fine-tuned by fine-tuning training dataset and evaluated on fine-tuning testing dataset
Used 50% PartMC data to train the pre-trained foundation model, fine-tuned by fine-tuning training dataset and evaluated on fine-tuning testing dataset
Used 90% PartMC data to train the pre-trained foundation model, fine-tuned by fine-tuning training dataset and evaluated on fine-tuning testing dataset
Fine-tuned foundation model ('Swapped Temporal Order' experiment), here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning testing dataset), model performance evaluated on fine-tuning training dataset
Model trained and selected by AutoML, here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning training dataset)
Model trained and selected by AutoML ('Swapped Temporal Order' experiment), here XX means training dataset is XX fraction of total MEGAPOLI data (XX *2 fraction of fine-tuning testing dataset), model performance evaluated on fine-tuning training dataset
This work made use of the facilities of the N8 Centre of Excellence in Computationally Intensive Research (N8 CIR) provided and funded by the N8 research partnership and EPSRC (Grant No. EP/T022167/1). The Centre is coordinated by the Universities of Durham, Manchester and York.
The authors acknowledge the assistance given by Research IT and Computational Shared Facility 3 (CSF3) at The University of Manchester.
Z.Z. appreciates the support provided by the academic start-up funds from the Department of Earth and Environmental Sciences at The University of Manchester and the Aerosol Science Career Development Grant from The Aerosol Society. The authors declare no conflict of interest.