Skip to content

Latest commit

 

History

History
196 lines (170 loc) · 7.88 KB

File metadata and controls

196 lines (170 loc) · 7.88 KB

nifty-manuscript

Contains the code and figures for the manuscript associated with NIFty: Classification with Missing Data - A NIFty Pipeline for Single-Cell Proteomics

Download or clone the repository before continuing.

Data Prep

To download and prepare the manuscript datasets for NIFty, do the following:

  1. Ensure you have the following python packages installed in your environment:
    • pandas
    • numpy
    • scikit-learn
    • anndata
    • pyarrow
  2. Ensure you have the following R packages installed in your environment:
    • MSstats
    • tidyverse
    • data.table
    • readxl
  3. Navigate to Data_Prep/Ai_Van-Eyk_2025/Original_Data/README.md and follow the instructions.
  4. Navigate to Data_Prep/Furtwaengler_Porse_Schoof_2025/Original_Data/README.md and follow the instructions.
  5. Navigate to Data_Prep/Khan_Elcheikhali_Slavov_2024/Original_Data/README.md and follow the instructions.
  6. Navigate to Data_Prep/Leduc_Slavov_2022/Original_Data/README.md and follow the instructions.
  7. Navigate to Data_Prep/Montalvo_Alvarez-Dominguez_Slavov_2023/Original_Data/README.md and follow the instructions.
  8. Navigate to Data_Prep/Petrosius_Schoof_2025/Original_Data/README.md and follow the instructions.
  9. Navigate to Data_Prep/Saddic_Parker_2026/Original_Data/README.md and follow the instructions.

Testing

If you have not already, follow the instructions in the Data Prep section before continuing.

Incomplete Data

To recreate the results for testing on incomplete data, do the following:

  1. Ensure you have the following python packages installed in your environment:

    • pandas
    • numpy
  2. Download NIFty and install the dependencies as described in the documentation.

  3. Navigate to Testing/No_Imputation/ and run the following python files:

    • Ai_imputed.py
    • Ai_unimputed.py
    • Furtwaengler_imputed_HSCxEarlyEryth.py
    • Furtwaengler_imputed_HSCxEMP.py
    • Furtwaengler_unimputed_HSCxEarlyEryth.py
    • Furtwaengler_unimputed_HSCxEMP.py
    • Khan_imputed.py
    • Khan_unimputed.py
    • Leduc_imputed.py
    • Leduc_unimputed.py
    • Montalvo_imputed.py
    • Montalvo_unimputed.py
    • Petrosius_imputed.py
    • Petrosius_unimputed.py
    • Saddic_imputed_fibro.py
    • Saddic_unimputed_fibro.py
    • Saddic_imputed_mfn.py
    • Saddic_unimputed_mfn.py
    • Saddic_imputed_smc.py
    • Saddic_unimputed_smc.py
    • Saddic_imputed_wt.py
    • Saddic_unimputed_wt.py
  4. For each configuration file created (found in Testing/No_Imputation/Test_{Dataset Identifier}/Config_Files), run NIFty using the following command:

    python <path_to_local_NIFty_download>/nifty.py -c <path to config file>

  5. Run combine_results.py (this will replace combined_results.tsv with your results).

Method Comparison

To recreate the results for comparing NIFty with traditional classification methods, do the following:

  1. Ensure you have the following python packages installed in your environment:

    • pandas
    • numpy
    • tomllib
    • sklearn
  2. Download NIFty and install the dependencies as described in the documentation.

  3. Navigate to Testing/Method_Comparison/Random_Forest_Tests/ and run the following python files:

    • Ai_imputed.py
    • Ai_unimputed.py
    • Furtwaengler_imputed_HSCxEarlyEryth.py
    • Furtwaengler_imputed_HSCxEMP.py
    • Furtwaengler_unimputed_HSCxEarlyEryth.py
    • Furtwaengler_unimputed_HSCxEMP.py
    • Khan_imputed.py
    • Khan_unimputed.py
    • Leduc_imputed.py
    • Leduc_unimputed.py
    • Montalvo_imputed.py
    • Montalvo_unimputed.py
    • Petrosius_imputed.py
    • Petrosius_unimputed.py
    • Saddic_imputed_fibro.py
    • Saddic_unimputed_fibro.py
    • Saddic_imputed_mfn.py
    • Saddic_unimputed_mfn.py
    • Saddic_imputed_smc.py
    • Saddic_unimputed_smc.py
    • Saddic_imputed_wt.py
    • Saddic_unimputed_wt.py
  4. For each configuration file created (found in Testing/Method_Comparison/Random_Forest_Tests/Test_{Dataset Identifier}/Config_Files), run the following command from the Testing/Method_Comparison/Random_Forest_Tests/ directory:

    python train_model.py <path to config file>

  5. Run combine_results.py (this will replace combined_results_random_forest.tsv with your results).

  6. Navigate to Testing/Method_Comparison/SVM_Tests/ and run the following python files:

    • Ai_imputed.py
    • Ai_unimputed.py
    • Furtwaengler_imputed_HSCxEarlyEryth.py
    • Furtwaengler_imputed_HSCxEMP.py
    • Furtwaengler_unimputed_HSCxEarlyEryth.py
    • Furtwaengler_unimputed_HSCxEMP.py
    • Khan_imputed.py
    • Khan_unimputed.py
    • Leduc_imputed.py
    • Leduc_unimputed.py
    • Montalvo_imputed.py
    • Montalvo_unimputed.py
    • Petrosius_imputed.py
    • Petrosius_unimputed.py
    • Saddic_imputed_fibro.py
    • Saddic_unimputed_fibro.py
    • Saddic_imputed_mfn.py
    • Saddic_unimputed_mfn.py
    • Saddic_imputed_smc.py
    • Saddic_unimputed_smc.py
    • Saddic_imputed_wt.py
    • Saddic_unimputed_wt.py
  7. For each configuration file created (found in Testing/Method_Comparison/SVM_Tests/Test_{Dataset Identifier}/Config_Files), run the following command from the Testing/Method_Comparison/SVM_Tests/ directory:

    python train_model.py <path to config file>

  8. Run combine_results.py (this will replace combined_results_svm.tsv with your results).

  9. Navigate to Testing/Method_Comparison/NIFty_SVM_Tests/ and run the following python files:

    • Ai_imputed.py
    • Ai_unimputed.py
    • Furtwaengler_imputed_HSCxEarlyEryth.py
    • Furtwaengler_imputed_HSCxEMP.py
    • Furtwaengler_unimputed_HSCxEarlyEryth.py
    • Furtwaengler_unimputed_HSCxEMP.py
    • Khan_imputed.py
    • Khan_unimputed.py
    • Leduc_imputed.py
    • Leduc_unimputed.py
    • Montalvo_imputed.py
    • Montalvo_unimputed.py
    • Petrosius_imputed.py
    • Petrosius_unimputed.py
    • Saddic_imputed_fibro.py
    • Saddic_unimputed_fibro.py
    • Saddic_imputed_mfn.py
    • Saddic_unimputed_mfn.py
    • Saddic_imputed_smc.py
    • Saddic_unimputed_smc.py
    • Saddic_imputed_wt.py
    • Saddic_unimputed_wt.py
  10. For each configuration file created (found in Testing/Method_Comparison/NIFty_SVM_Tests/Test_{Dataset Identifier}/Config_Files), run the following command:

    python <path_to_local_NIFty_download>/nifty.py -c <path to config file>

  11. Run combine_results.py (this will replace combined_results_NIFty_SVM.tsv with your results).

Batch Effects

Upcoming

Multiclass

To recreate the results for testing on multiclass data, do the following:

  1. Ensure you have the following python packages installed in your environment:
    • pandas
    • scikit-learn
  2. Download NIFty from and install the dependencies as described in the documentation.
  3. Navigate to Testing/Multiclass/.
  4. Add the absolute path to the directory containing nifty.py and associated files to line 18 in multiclass_wrapper.py.
  5. Run multiclass_wrapper.py.
  6. Run combine_predictions.py (this will replace combined_predictions.tsv with your results).

Figures and Tables

You do not need to have completed the Data Prep and Testing sections before continuing.

To recreate Figures 3, 4, and 6 and Table 1 found in the manuscript, do the following:

  1. Ensure you have the following R packages installed in your environment:
    • tidyverse
    • this.path
    • caret
    • ggtext
  2. Navigate to Figures_and_Tables.
  3. Run the following R files:
    • generate_Figure_3.R (recreates Fig3_Leduc.png, Fig3_Montalvo.png)
    • generate_Table_1.R (recreates Table1.tsv)
    • generate_Figure_4.R (recreates Fig4.png)
    • generate_Figure_5.R (recreates Fig5.png)
    • generate_Figure_7.R (recreates Fig7.png)