Automated ischemic stroke lesion segmentation in native-space T1-weighted MRI using the ATLAS v3.0 dataset (N=1453). Submission to ISLES 2026 Grand Challenge @ MICCAI 2026.
This project implements a 3D U-Net with pluggable metadata conditioning for stroke lesion segmentation. Two interchangeable tracks:
| Track | Method | Speed | Quality |
|---|---|---|---|
| A | FiLM (Feature-wise Linear Modulation) | Fast | Strong baseline |
| C | LLM Conditioning (all-MiniLM-L6-v2) | Slower | Potentially higher quality |
# PyTorch (CUDA 12.1)
pip install torch==2.3.1 torchvision==0.18.1 --index-url https://download.pytorch.org/whl/cu121
# Other dependencies
pip install -r requirements.txtBefore running any pipeline, configure your data paths in configs/config.yaml:
data:
root: ./data/raw/ATLAS3_Training_Raw # Update to your dataset location
processed_dir: ./data/processed # Update to your output location
logging:
log_dir: ./outputs/logs # Update to your log locationPath configuration guidelines:
- Use relative paths (e.g.,
./data/...) for portability - On Windows: use forward slashes or escaped backslashes (
C:/data/...orC:\\data\\...) - On Linux/Mac: use standard Unix paths (
/home/user/data/...)
Or override paths via command line:
python pipeline/preprocessing.py \
--config configs/config.yaml \
data.root=/path/to/data \
data.processed_dir=/path/to/outputThe dataset must be preprocessed before training:
# Preprocess raw NIfTI files
python pipeline/preprocessing.py \
--config configs/config.yaml \
--workers 4
# Generate 5-fold CV splits
python pipeline/splits.py \
--config configs/config.yaml \
--inspect# Train a single fold (Track A)
python pipeline/train.py \
--config configs/config.yaml \
--fold 0 \
--track A \
--model-size small # tiny or base
# Train all 5 folds
python pipeline/train.py \
--config configs/config.yaml \
--fold all \
--track A# Evaluate with Test-Time Augmentation
python pipeline/evaluate.py \
--config configs/config.yamlRun overnight training, automated evaluation, OOM-resilient execution, and Docker/CPU submission verification across all model sizes (tiny, small, base) and tracks (A, C):
# 1. Quick comparative benchmark across all sizes & tracks (Fold 0):
python scripts/master_benchmark.py --config configs/config.yaml --fold 0
# 2. Full 5-fold CV overnight training + automatic evaluation across all models:
python scripts/master_benchmark.py --config configs/config.yaml --fold all --mode all
# 3. Evaluate all ready models:
python scripts/master_benchmark.py --config configs/config.yaml --mode eval
# 4. Verify Docker / CPU submission latency & shape constraints:
python scripts/master_benchmark.py --config configs/config.yaml --mode verify
# 5. Helper script run with auto multi-GPU DataParallel:
python scripts/master_benchmark.py --config configs/config.yaml --mode all --name overnight_all_modelsThe master benchmark script includes automatic GPU detection and load balancing:
# Run with auto GPU selection (DataParallel on available GPUs):
python scripts/master_benchmark.py --config configs/config.yaml --fold all --mode all
# Monitor GPU usage
nvidia-smi -l 5Job types:
master— Master overnight runner (all models, automated eval, OOM resilience, Docker check)preprocess— Run preprocessing pipelinesplits— Generate CV splitstrain— Train model (add--fold allfor all 5 folds)evaluate— Evaluate with TTA
Options:
--name <job_name>— Custom log file name--fold <n|all>— Which fold(s) to train--track <A|C>— Which conditioning track--workers <n>— Number of parallel workers--tta— Enable test-time augmentation--local— Run without nohup (for testing)
Package, build, and export the Docker container for Grand Challenge submission:
# 1. Package best available model (auto-copies checkpoints + config):
python scripts/package_submission.py --track A --size tiny
# 2. Package AND build Docker image + export .tar.gz:
python scripts/package_submission.py --track A --size tiny --build --export
# 3. Package all 5 folds for ensemble submission:
python scripts/package_submission.py --track A --size base --folds all --build --export
# 4. Auto-package best model from master benchmark pipeline:
python scripts/master_benchmark.py --config configs/config.yaml --mode package
# 5. Full pipeline: train + eval + verify + auto-package + build Docker:
python scripts/master_benchmark.py --config configs/config.yaml --mode all --build-dockerWhat package_submission.py does:
- Finds trained
best.pthcheckpoints for the specified track/size/folds - Copies them into
checkpoints/fold_0_best.pth,fold_1_best.pth, etc. - Saves a submission config with the correct track and model size
- Builds Docker image:
docker build -t isles26-submission . - Exports
.tar.gzcontainer:docker save | gzip > isles26_submission.tar.gz
film-isles26/
├── configs/ # Configuration files
│ ├── config.yaml # Master config (all modules read from here)
│ ├── config_rtx.yaml # RTX workstation config (deprecated)
│ ├── track_A.yaml # Track A override (FiLM conditioning)
│ └── track_C.yaml # Track C override (LLM conditioning)
├── pipeline/
│ ├── preprocessing.py # Reorientation, clipping, z-score normalization
│ ├── splits.py # 5-fold CV, stratified by CHRONICITY_DERIVED × SITE
│ ├── augmentation.py # MONAI transforms + phase-specific augmentation
│ ├── dataset.py # PyTorch Dataset, metadata encoding, DataLoader
│ ├── conditioning.py # FiLMConditioner (A) + LLMConditioner (C)
│ ├── model.py # 3D U-Net + FiLM injection + deep supervision
│ ├── loss.py # Dice + CE + boundary focal loss
│ ├── train.py # Poly LR, AdamW, mixed precision, early stopping
│ ├── evaluate.py # All 5 official metrics
│ ├── visualize.py # Training curves, overlays, track comparison
│ └── tests/
│ └── test_pipeline.py # Unit tests
├── utils/
│ ├── __init__.py
│ └── eval_utils.py # ISLES26 official evaluation metrics
├── scripts/
│ ├── master_benchmark.py # Multi-model orchestration & benchmarking
│ ├── package_submission.py # Docker submission packaging
│ └── interpretability.py # Post-hoc analysis (local only)
├── entrypoint.py # Docker entrypoint for Grand Challenge API
├── notebooks/
│ ├── ingest_atlas.ipynb # Data ingestion from Kaggle
│ ├── eda-atlas.ipynb # Exploratory data analysis
│ └── smoke_tests.ipynb # Kaggle-compatible pytest runner
├── checkpoints/ # Trained model weights (generated at runtime)
├── outputs/ # Training logs and results (generated at runtime)
├── data/ # Raw and processed data (generated at runtime)
├── Dockerfile # Docker image definition
├── README.md
├── CLAUDE.md # Project context for AI assistant
├── requirements.txt
└── pytest.ini
ATLAS v3.0 (ISLES26 training set):
- N = 1453 sessions from 33 sites (R001–R052 + SOOP)
- Native-space skull-stripped T1w MRI
- Single lesion mask per session
- Metadata: DAYS_POST_STROKE, CHRONICITY (1 or NaN), CHRONICITY_DERIVED, SITE
Key characteristics:
- Orientations: RAS (60%), LAS (40%) — reorientation mandatory
- Spacing: isotropic ~1mm³ — no resampling needed
- Inter-site intensity CV: 1.453 — per-scan z-score required
- Lesion sizes: small <1mL (27%), medium 1-10mL (37%), large >10mL (37%)
5-dimensional metadata vector:
[days_norm, is_acute, is_subacute, is_chronic, confirmed_chronic]
days_norm: log1p(days) / log1p(10000) in [0, 1]is_acute/subacute/chronic: One-hot encoding from CHRONICITY_DERIVEDconfirmed_chronic: 1.0 if CHRONICITY == 1.0 (organizer-provided)
Natural language string:
"Stroke MRI scan from site R001. Time since stroke: 45 days. Phase: chronic. Task: segment the ischemic lesion."
- Dice Score — global binary DSC
- Absolute Volume Difference — mL
- Absolute Lesion Count Difference — instance count |GT − Pred|
- Lesion-wise F1 — recognition quality (panoptica, threshold=0.25)
- PR-AUC — requires soft probability map (not binary mask)
- Raw data: ~8GB (compressed) / ~16GB (extracted)
- Preprocessed: ~10-12GB
- Training: Additional ~5GB for checkpoints per model configuration
- Time limit: 10 minutes per scan on T4 GPU
- Memory: 32GB RAM
- Native space only: No registration to MNI allowed in final output
- Track A only: Track C adds ~2-3 min LLM loading overhead
Track C (LLM) is a time risk — benchmark Track A first.
This project is designed for Grand Challenge Docker submission. The pipeline:
- Loads a single T1w NIfTI input
- Reorients to RAS, clips, and normalizes
- Runs ensemble prediction across 5 fold models
- Applies thresholding and connected component removal
- Reorients output back to input orientation
- Docker installed (version 20.10+ recommended)
- Trained model checkpoints (
fold_0_best.pththroughfold_4_best.pth) - Test NIfTI scan (can be any ATLAS training scan)
# From the project root (where Dockerfile is located)
docker build -t isles26-submission .The Docker image expects checkpoints in /opt/algorithm/checkpoints/:
mkdir -p checkpoints
cp path/to/fold_0_best.pth checkpoints/
cp path/to/fold_1_best.pth checkpoints/
cp path/to/fold_2_best.pth checkpoints/
cp path/to/fold_3_best.pth checkpoints/
cp path/to/fold_4_best.pth checkpoints/Option A: Copy checkpoints into image (recommended for final submission)
# Dockerfile line 34 already includes: COPY checkpoints/ /opt/algorithm/checkpoints/
docker build -t isles26-submission .Option B: Mount checkpoints at runtime (for testing)
docker build -t isles26-submission .
docker run --gpus all \
-v /path/to/checkpoints:/opt/algorithm/checkpoints \
isles26-submission# Test with a local scan
docker run --gpus all \
-v /path/to/test_input:/input \
-v /path/to/test_output:/output \
isles26-submissionpython -c "
import nibabel as nib
import numpy as np
inp = nib.load('/path/to/test_input/image.nii.gz')
out = nib.load('/path/to/test_output/mask.nii.gz')
assert inp.shape == out.shape, f'Shape mismatch: {inp.shape} vs {out.shape}'
assert np.allclose(inp.affine, out.affine), 'Affine mismatch'
data = out.get_fdata()
assert set(np.unique(data)).issubset({0.0, 1.0}), 'Non-binary values found'
print('Geometry check passed!')
"To test with the Grand Challenge input format:
# Create test input directory structure
mkdir -p /tmp/isles_test/input/images/t1-brain-mri
cp /path/to/test_scan.nii.gz /tmp/isles_test/input/images/t1-brain-mri/
# Run with Docker
docker run --rm \
-v /tmp/isles_test/input:/input \
-v /tmp/isles_test/output:/output \
-v /path/to/checkpoints:/opt/algorithm/checkpoints \
isles26-submissionKey files in the Docker image:
/opt/algorithm/entrypoint.py— Main inference script/opt/algorithm/pipeline/— Pipeline modules/opt/algorithm/utils/— Evaluation utilities/opt/algorithm/checkpoints/— 5 fold models
- Output mask geometry matches input exactly (shape + affine)
- Container runs in < 10 min on a single scan on T4
- No internet access required inside container (all weights bundled)
- Empty mask output works correctly for healthy scans
- Test on both RAS and LAS input orientations
- Test on small and large lesion cases
- Submit to preliminary phase first (2 debug scans)
| Decision | Choice | Reason |
|---|---|---|
| Backbone | Custom 3D U-Net | Full control over conditioning injection |
| Conditioning | Decoder bottleneck | Full receptive field before modulation |
| Normalization | Per-scan foreground z-score | Inter-site CV=1.453 |
| Augmentation | MONAI + phase-specific | Acute: blur; Chronic: cavity inversion |
| Loss | Dice + CE + boundary focal | Class imbalance + small lesion upweighting |
| Training | 5-fold CV, poly LR, AdamW | Maximize use of 1453 scans |
-
Memory error during preprocessing
- Reduce
num_workersin preprocessing config - Use
/kaggle/tempinstead of/kaggle/workingon Kaggle
- Reduce
-
CUDA out of memory
- Reduce
training.batch_size - Enable
training.mixed_precision
- Reduce
-
Missing CHRONICITY column
- Re-run preprocessing with latest metadata
- Check
metadata/metadata.csvexists
-
Config path errors
- Verify paths in
configs/config.yamlexist - Use forward slashes on Windows:
C:/path/to/data
- Verify paths in
| Section | Key | Default | Description |
|---|---|---|---|
data.root |
- | ./data/raw/ATLAS3_Training_Raw |
Path to raw dataset directory |
data.processed_dir |
- | ./data/processed |
Path for preprocessed data output |
logging.log_dir |
- | ./outputs/logs |
Path for training logs and checkpoints |
preprocessing.target_orientation |
- | RAS |
Target orientation for all scans |
preprocessing.normalization |
- | per_scan_zscore |
Normalization method |
training.epochs |
- | 1000 |
Maximum training epochs |
training.batch_size |
- | 4 |
Batch size per GPU |
conditioning.track |
- | A |
Conditioning track: A (FiLM) or C (LLM) |
model.size |
- | tiny |
Model size: tiny, small, or base |
| Variable | Description |
|---|---|
PYTHONUNBUFFERED=1 |
Force Python output to be unbuffered |
PYTHONDONTWRITEBYTECODE=1 |
Prevent .pyc file generation in Docker |
| Phase | Status | Description |
|---|---|---|
| Phase 1 | Completed | Verification, smoke tests, data ingestion |
| Phase 2 | In progress | Baseline training (Track A) |
| Phase 3 | Pending | Track C experiment |
| Phase 4 | Pending | Track B pseudo-labelling |
| Phase 5 | Pending | Docker submission |
| Phase 6 | Pending | Paper writing |
Joseph Derrick
GHAiC-K Lab, Kumasi, Ghana
MIT License