Skip to content

About

Classifying anemia type from full blood count parameters

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

11 Commits

Folders and files

Repository files navigation

Anemia Classification from Full Blood Count Parameters

Problem statement

Anemia is a reduction in circulating haemoglobin below the reference range for a patient's age and sex. It is not a single disease: iron deficiency anemia, megaloblastic anemia (B12/folate deficiency), and anemia of chronic disease (ACD) arise from distinct mechanisms, present with overlapping full blood count (FBC) pictures at a glance, and are managed differently. A haematologist narrows between them using a small set of red cell indices before ordering confirmatory iron studies, B12, folate, or inflammatory markers. This project asks whether a classifier trained on the eight parameters that any automated FBC analyser already reports — haemoglobin (Hb), red cell count (RBC), mean corpuscular volume (MCV), mean corpuscular haemoglobin (MCH), mean corpuscular haemoglobin concentration (MCHC), red cell distribution width (RDW), platelet count, and white cell count (WBC) — can reproduce that first-pass triage, and whether its reasoning matches the morphological logic a haematologist would use.

Data

No public dataset combining these eight FBC parameters with a clean, four-way anemia-type label (normal / iron deficiency / megaloblastic / anemia of chronic disease) was reachable. The cohort in this repository is synthetic, generated by src/data.py, and does not represent any real patient.

Generation logic, per class:

  • Normal: all eight parameters drawn around adult reference-range means.
  • Iron deficiency anemia: microcytic (low MCV), hypochromic, and critically, a raised RDW — the anisocytosis produced by a mixed population of older normocytic and newly formed microcytic red cells. Reactive thrombocytosis is included, as is commonly seen.
  • Megaloblastic anemia: macrocytic (high MCV), RDW raised but less than in iron deficiency, with mild leukopenia and thrombocytopenia reflecting the marrow maturation defect.
  • Anemia of chronic disease: normocytic and normochromic with a normal RDW — the feature that separates it from iron deficiency despite a similar MCV — and a raised WBC reflecting the underlying inflammatory, infective, or malignant process driving the anemia.

RBC, MCV and MCHC are sampled per class first; Hct is derived as RBC * MCV / 10 and Hb as MCHC * Hct / 100 plus small measurement noise, so that the eight reported values are internally consistent in the way a real analyser's outputs are, rather than independently sampled. Class sizes (1300 patients total: 400 normal, 450 iron deficiency, 300 chronic disease, 150 megaloblastic) reflect roughly the relative frequency of these presentations, with megaloblastic anemia the rarest. Full parameter choices and the code that implements this logic are in src/data.py.

Methods

  • Feature engineering (src/features.py): the eight raw parameters plus two derived indices — Mentzer index (MCV / RBC), used at the bench to help separate iron deficiency from other microcytic pictures, and an RDW x MCV interaction term that gives models direct access to the "low MCV, high RDW" signature without requiring them to discover it.
  • Models (src/model.py): logistic regression (with standardised inputs) as an interpretable baseline, compared against random forest and gradient boosting. Both ensembles are grid-searched (5-fold, macro F1) in the modeling notebook; tuning brings gradient boosting to parity with logistic regression's untuned score and random forest to just under it, so logistic regression is kept as the reported model on parsimony and interpretability grounds.
  • Validation: stratified 5-fold cross-validation on an 80% training split for model selection, out-of-fold predictions used for every cross-validation metric reported. Final numbers are from a held-out 20% test split none of the models saw during training or selection.
  • Interpretation: permutation importance (F1-macro scoring) for both the reported logistic regression and the random forest, plus a SHAP summary for the random forest, checked against the MCV/RDW logic used to generate the labels.

Results

Five-fold cross-validation, macro-averaged (training split, out-of-fold):

model precision recall f1
logistic regression 0.955 0.949 0.952
random forest 0.948 0.943 0.945
gradient boosting 0.946 0.941 0.944

Held-out test set (logistic regression, the reported model — chosen on parsimony grounds since it does not trail the ensembles):

class precision recall f1 support
normal 0.951 0.975 0.963 80
iron deficiency 0.989 1.000 0.994 90
megaloblastic 0.938 1.000 0.968 30
chronic disease 0.964 0.883 0.922 60
macro avg 0.960 0.965 0.962 260

Figures (reports/figures/):

  • param_distributions_by_class.png — violin plots of all eight parameters by class
  • mcv_vs_rdw.png — MCV vs RDW scatter, coloured by class
  • confusion_matrix_test.png — test-set confusion matrix
  • shap_summary.png — SHAP feature importance, random forest

Clinical interpretation

Chronic disease is where the model is weakest (88.3% recall): its false negatives are almost all misclassified as normal or iron deficiency, which is exactly where a haematologist also struggles — ACD can present with a near-normal FBC when the underlying disease is mild, and a mildly raised WBC alone is a weak discriminator against a merely elevated-but-normal count. This is a genuine limitation to flag before deploying anything like this as a triage aid: missing ACD does not usually delay urgent treatment (the anemia itself is rarely severe), but it does mean the underlying inflammatory or malignant driver goes uninvestigated.

Permutation importance shows hb dominating for both models, which follows from how the data was generated: Hb is derived from RBC, MCV and MCHC together, so it carries a blended read of severity across all three. Beyond that, logistic regression retains real weight on rdw, rbc, mcv, and the Mentzer index, tracking the morphological rules used to build the labels. The random forest instead concentrates importance on hb and mch, assigning little individually to mcv and rdw — most likely because permutation importance splits credit across the correlated raw and engineered features (mcv also appears inside mentzer_index and rdw_mcv_interaction), not because the forest ignores morphology; its confusion matrix shows it separating classes on the same pattern. A model that only ever reports hb as important would be clinically uninformative on its own — Hb by itself does not distinguish anemia type — so the secondary ranking on RDW and MCV is the part worth checking, and cross-checking importance across more than one model turned out to matter here.

Limitations

  • The cohort is synthetic. Class boundaries were generated from independent per-class parameter distributions with fixed correlation structure (via the Hb/Hct derivation only), so the classes are more separable here than in a real, noisier population with comorbidities, combined deficiencies, or partially treated anemia.
  • No mixed or transitional cases are modelled (e.g. combined iron and B12 deficiency, or early/partially treated ACD), which are common in practice and are exactly where a real classifier would be tested hardest.
  • Thalassemia trait, which overlaps with iron deficiency on MCV and Mentzer index, is not a modelled class; a real deployment would need to include it or explicitly state the exclusion.
  • Reference intervals used are adult, unisex approximations; real FBC interpretation is age- and sex-stratified.

How to reproduce

python -m venv venv
source venv/Scripts/activate  # venv\Scripts\activate on cmd.exe
pip install -r requirements.txt

python src/data.py                                        # generate data/raw/cbc_cohort.csv
jupyter nbconvert --to notebook --execute --inplace notebooks/01_eda.ipynb
jupyter nbconvert --to notebook --execute --inplace notebooks/02_modeling.ipynb
pytest tests/

About

Classifying anemia type from full blood count parameters

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages