Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

A Simplified Drug-Response Model, with Dropout Intervals

The comparison baseline: a lighter architecture that produces uncertainty by Monte Carlo Dropout.

Ved Piyush · PhD in Statistics · University of Nebraska–Lincoln


Dropout randomly switches off parts of a neural network during training to prevent it from memorising. Monte Carlo Dropout is a trick: leave dropout switched on at prediction time too, run the same input through many times, and read the spread of answers as an uncertainty estimate. It is the standard, easy baseline that any new uncertainty method has to beat.

This repository holds a pared-down version of the drug-response architecture together with that baseline, so the comparison against the ensemble Kalman filter method is like-for-like on identical data.

How it fits together

flowchart LR
    A["Drug + cell-line<br/>features"] --> B["Simplified<br/>network"]
    B --> C["Dropout left on<br/>at prediction"]
    C --> D["Many forward<br/>passes"]
    D --> E["Interval<br/>+ coverage"]
    style C fill:#0b7a64,color:#fff,stroke:#0b7a64
Loading

What the code does

SimpleCDRGCN_Dropout_Intervals produces Monte Carlo Dropout intervals and computes coverage. Other notebooks vary the pieces that matter for a fair comparison:

  • dropout kept active in training only, versus in training and prediction (..._active_only_train_not_pred versus ..._active_both);
  • a no-leakage variant ensuring no cell line appears in both training and test (..._no_leakage), and a transfer-learning version of it;
  • ten-fold cross-validation built from held-out samples;
  • feature construction from drug structure and cell-line molecular data, with and without normalisation.

Uses TensorFlow Probability for the probabilistic layers and RDKit/DeepChem to turn drug structures into graphs.

Repository layout

SimplerDeepCDR/ holds the notebooks behind the reported results; SimplerDeepCDR/Dev_Scripts/ holds earlier exploratory versions, including alternative inputs such as 1-D convolutions over chemical strings.

Running it

  1. SimplerDeepCDR_Drug_Cell_Line_Features_Create_Feature_Mats.ipynb builds the feature matrices.
  2. SimplerCDR_Exact_Network_more_dropout_no_leakage.ipynb trains without train/test leakage.
  3. SimpleCDRGCN_Dropout_Intervals.ipynb produces the Monte Carlo Dropout intervals and coverage.

Where to look first

  • SimplerDeepCDR/SimpleCDRGCN_Dropout_Intervals.ipynb — the Monte Carlo Dropout intervals and their coverage
  • SimplerDeepCDR/SimplerCDR_Exact_Network_more_dropout_no_leakage.ipynb — the leakage-free comparison run

What is in each directory

Directory Files Purpose
SimplerDeepCDR 12 notebooks, 1 script the reported baseline runs
SimplerDeepCDR/Dev_Scripts exploratory 19 notebooks, 4 scripts exploratory versions kept for provenance — not needed to reproduce results

Directories marked exploratory are earlier iterations kept for provenance. To reproduce the reported results you need only the 1 core directory above.

Notes

Notebook outputs are committed, so the figures and result tables render on GitHub without running anything. Molecular feature matrices and trained weights are not committed; the feature notebooks rebuild them.

Research code from my doctoral work at the University of Nebraska–Lincoln (31 notebooks). Previously hosted at github.com/Ved-Piyush/DeepCDR_SimpleCDR.


Ved Piyush, PhD · Website · Google Scholar · vedpiyush93@gmail.com

About

A simplified cancer drug-response model with Monte Carlo Dropout prediction intervals - the baseline the ensemble Kalman filter method is measured against.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages