Anemia is a reduction in circulating haemoglobin below the reference range for a patient's age and sex. It is not a single disease: iron deficiency anemia, megaloblastic anemia (B12/folate deficiency), and anemia of chronic disease (ACD) arise from distinct mechanisms, present with overlapping full blood count (FBC) pictures at a glance, and are managed differently. A haematologist narrows between them using a small set of red cell indices before ordering confirmatory iron studies, B12, folate, or inflammatory markers. This project asks whether a classifier trained on the eight parameters that any automated FBC analyser already reports — haemoglobin (Hb), red cell count (RBC), mean corpuscular volume (MCV), mean corpuscular haemoglobin (MCH), mean corpuscular haemoglobin concentration (MCHC), red cell distribution width (RDW), platelet count, and white cell count (WBC) — can reproduce that first-pass triage, and whether its reasoning matches the morphological logic a haematologist would use.
No public dataset combining these eight FBC parameters with a clean,
four-way anemia-type label (normal / iron deficiency / megaloblastic /
anemia of chronic disease) was reachable. The cohort in this repository
is synthetic, generated by src/data.py, and does not represent any
real patient.
Generation logic, per class:
- Normal: all eight parameters drawn around adult reference-range means.
- Iron deficiency anemia: microcytic (low MCV), hypochromic, and critically, a raised RDW — the anisocytosis produced by a mixed population of older normocytic and newly formed microcytic red cells. Reactive thrombocytosis is included, as is commonly seen.
- Megaloblastic anemia: macrocytic (high MCV), RDW raised but less than in iron deficiency, with mild leukopenia and thrombocytopenia reflecting the marrow maturation defect.
- Anemia of chronic disease: normocytic and normochromic with a normal RDW — the feature that separates it from iron deficiency despite a similar MCV — and a raised WBC reflecting the underlying inflammatory, infective, or malignant process driving the anemia.
RBC, MCV and MCHC are sampled per class first; Hct is derived as
RBC * MCV / 10 and Hb as MCHC * Hct / 100 plus small measurement
noise, so that the eight reported values are internally consistent in
the way a real analyser's outputs are, rather than independently
sampled. Class sizes (1300 patients total: 400 normal, 450 iron
deficiency, 300 chronic disease, 150 megaloblastic) reflect roughly the
relative frequency of these presentations, with megaloblastic anemia the
rarest. Full parameter choices and the code that implements this logic
are in src/data.py.
- Feature engineering (
src/features.py): the eight raw parameters plus two derived indices — Mentzer index (MCV / RBC), used at the bench to help separate iron deficiency from other microcytic pictures, and an RDW x MCV interaction term that gives models direct access to the "low MCV, high RDW" signature without requiring them to discover it. - Models (
src/model.py): logistic regression (with standardised inputs) as an interpretable baseline, compared against random forest and gradient boosting. Both ensembles are grid-searched (5-fold, macro F1) in the modeling notebook; tuning brings gradient boosting to parity with logistic regression's untuned score and random forest to just under it, so logistic regression is kept as the reported model on parsimony and interpretability grounds. - Validation: stratified 5-fold cross-validation on an 80% training split for model selection, out-of-fold predictions used for every cross-validation metric reported. Final numbers are from a held-out 20% test split none of the models saw during training or selection.
- Interpretation: permutation importance (F1-macro scoring) for both the reported logistic regression and the random forest, plus a SHAP summary for the random forest, checked against the MCV/RDW logic used to generate the labels.
Five-fold cross-validation, macro-averaged (training split, out-of-fold):
| model | precision | recall | f1 |
|---|---|---|---|
| logistic regression | 0.955 | 0.949 | 0.952 |
| random forest | 0.948 | 0.943 | 0.945 |
| gradient boosting | 0.946 | 0.941 | 0.944 |
Held-out test set (logistic regression, the reported model — chosen on parsimony grounds since it does not trail the ensembles):
| class | precision | recall | f1 | support |
|---|---|---|---|---|
| normal | 0.951 | 0.975 | 0.963 | 80 |
| iron deficiency | 0.989 | 1.000 | 0.994 | 90 |
| megaloblastic | 0.938 | 1.000 | 0.968 | 30 |
| chronic disease | 0.964 | 0.883 | 0.922 | 60 |
| macro avg | 0.960 | 0.965 | 0.962 | 260 |
Figures (reports/figures/):
param_distributions_by_class.png— violin plots of all eight parameters by classmcv_vs_rdw.png— MCV vs RDW scatter, coloured by classconfusion_matrix_test.png— test-set confusion matrixshap_summary.png— SHAP feature importance, random forest
Chronic disease is where the model is weakest (88.3% recall): its false negatives are almost all misclassified as normal or iron deficiency, which is exactly where a haematologist also struggles — ACD can present with a near-normal FBC when the underlying disease is mild, and a mildly raised WBC alone is a weak discriminator against a merely elevated-but-normal count. This is a genuine limitation to flag before deploying anything like this as a triage aid: missing ACD does not usually delay urgent treatment (the anemia itself is rarely severe), but it does mean the underlying inflammatory or malignant driver goes uninvestigated.
Permutation importance shows hb dominating for both models, which
follows from how the data was generated: Hb is derived from RBC, MCV
and MCHC together, so it carries a blended read of severity across all
three. Beyond that, logistic regression retains real weight on rdw,
rbc, mcv, and the Mentzer index, tracking the morphological rules
used to build the labels. The random forest instead concentrates
importance on hb and mch, assigning little individually to mcv
and rdw — most likely because permutation importance splits credit
across the correlated raw and engineered features (mcv also appears
inside mentzer_index and rdw_mcv_interaction), not because the
forest ignores morphology; its confusion matrix shows it separating
classes on the same pattern. A model that only ever reports hb as
important would be clinically uninformative on its own — Hb by itself
does not distinguish anemia type — so the secondary ranking on RDW and
MCV is the part worth checking, and cross-checking importance across
more than one model turned out to matter here.
- The cohort is synthetic. Class boundaries were generated from independent per-class parameter distributions with fixed correlation structure (via the Hb/Hct derivation only), so the classes are more separable here than in a real, noisier population with comorbidities, combined deficiencies, or partially treated anemia.
- No mixed or transitional cases are modelled (e.g. combined iron and B12 deficiency, or early/partially treated ACD), which are common in practice and are exactly where a real classifier would be tested hardest.
- Thalassemia trait, which overlaps with iron deficiency on MCV and Mentzer index, is not a modelled class; a real deployment would need to include it or explicitly state the exclusion.
- Reference intervals used are adult, unisex approximations; real FBC interpretation is age- and sex-stratified.
python -m venv venv
source venv/Scripts/activate # venv\Scripts\activate on cmd.exe
pip install -r requirements.txt
python src/data.py # generate data/raw/cbc_cohort.csv
jupyter nbconvert --to notebook --execute --inplace notebooks/01_eda.ipynb
jupyter nbconvert --to notebook --execute --inplace notebooks/02_modeling.ipynb
pytest tests/