Code and resources to support the main bootstrap-based L1-logistic-regression analysis used in this study.
This repository contains scripts and example data for the bootstrap-based feature selection analysis described in the paper.
Its purpose is to improve transparency and reproducibility of the main computational results.
Runs a single bootstrap iteration of L1-regularized logistic regression.
This step:
- resamples the input data with replacement,
- adds small Gaussian noise,
- fits an L1-penalized logistic regression model,
- returns the selected-feature mask (
|coef| > epsilon), coefficients, and out-of-bag (OOB) sample indices.
Runs repeated bootstrap iterations in parallel.
This step:
- performs
n_bootstrapiterations, - collects coefficients across iterations,
- returns a coefficient matrix of shape
(n_bootstrap × n_features)and the OOB index list.
Summarizes bootstrap results across iterations.
This step:
- computes feature-selection frequency,
- calculates coefficient means and standard deviations,
- identifies features passing a user-defined selection threshold,
- returns summary statistics for downstream interpretation.
Runs the main synthetic-data example from start to finish.
This script:
- loads the example input files,
- runs bootstrap-based L1-logistic regression,
- summarizes feature stability across bootstrap iterations,
- writes analysis outputs to the
results/directory.
The core workflow is:
- Prepare a processed feature matrix
Xand binary label vectory. - Run bootstrap-based L1-logistic regression across repeated resamples.
- Summarize selection frequency and coefficient stability across bootstrap iterations.
- Identify robustly selected features.
Clone the repository and create a clean virtual environment:
git clone https://github.com/kenflab/aml-venetoclax-resistance
cd aml-venetoclax-resistance
python3 -m venv .venv
source .venv/bin/activate # macOS / Linux
# .venv\Scripts\activate # Windows
pip install --upgrade pip
pip install -r requirements.txtTo run the main synthetic-data analysis:
python reproduce_main_analysis.pyThis command writes the following files to the results/ directory:
- bootstrap_coefficients.csv
- bootstrap_feature_summary.csv
- selected_features.json
- analysis_metadata.json
gene_names = X.columns
n_bootstrap = 10000
epsilon = 0.01
random_state = 2025
n_jobs = -1
threshold_ratio = 0.8
params = {
"C": 10,
"class_weight": "balanced",
"max_iter": 10000,
}
coef_matrix, oob_indices = fit_lasso_logistic_bootstrap(
X.values,
y,
gene_names=gene_names,
n_bootstrap=n_bootstrap,
epsilon=epsilon,
n_jobs=n_jobs,
random_state=random_state,
params=params,
)
summary = summarize_bootstrap_coefficients(
coef_matrix=coef_matrix,
gene_names=gene_names,
threshold_ratio=threshold_ratio,
)-
X: A numeric matrix where
- Rows represent samples
- Columns represent features (genes)
- Values are Transcripts Per Million (TPM), then log-transformed, and Z-score standardized per gene across samples.
-
y: A binary list or array of labels (
0and1), corresponding to the sample classes.
This repository includes synthetic example data for demonstration and code verification only.
-
The example data were generated using NumPy random numbers with a fixed seed (42).
-
No real patient or biological data are included in this repository.
-
Files are located in the
data/folder:data/X_sample.csv— Processed feature matrix (rows = samples, columns = genes).
Values are TPM-normalized, then log-transformed, and Z-score standardized per gene across samples.data/y_sample.csv— Labels (sample_id,label) with binary classes0/1.
Shapes
X_sample.csv: 30 × 50y_sample.csv: 30 × 2 (sample_id,label)
import pandas as pd
# Load X (processed) with sample IDs as index
X = pd.read_csv("data/X_sample.csv", index_col=0)
# Load y and align to X.index
y_df = pd.read_csv("data/y_sample.csv") # columns: sample_id, label
y = y_df.set_index("sample_id").loc[X.index, "label"].values
print(X.shape, y.shape)
print(X.head(3))
print(y[:10])This repository is intended to support the analysis presented in the manuscript. It is not packaged as a general-purpose software tool, but rather as a focused, minimal resource for reproducing the main bootstrap-based feature selection workflow.
For the original study data, please refer to the manuscript and its data availability statement.
