MSc Health Data Science candidate (URV, 2025 to 2027). MSc Molecular Biomedicine (Münster). Background in translational oncology.
I work at the point where clinical and commercial data turn into decisions: cohort studies on real electronic health records, statistical modelling, and dashboards that non technical stakeholders actually use.
| Project | What it shows |
|---|---|
| ICU lactate and mortality, interpreted by an LLM | End to end pipeline: SQL cohort extraction from MIMIC-III, multivariable logistic regression on 22,302 ICU stays, then an LLM that turns the model output into a publication style results paragraph. Built to stay inside the PhysioNet Data Use Agreement: aggregate statistics only. |
| Extracting heart failure signals from clinical notes | Regex based extraction of ejection fraction and medications from free text discharge summaries in MIMIC-III. Recovers the known bimodal HFrEF and HFpEF distribution, and documents where the extraction breaks rather than only where it works. |
| Heart disease markers, from analysis to dashboard | Exploratory analysis in Python, then an explanatory Power BI dashboard built for a clinical audience. Translating a statistical result into something a decision maker reads in thirty seconds. |
| When a classifier learns the speaker, not the disease | The UCI Parkinsons dataset holds 195 voice recordings from only 32 people. Splitting them at random puts the same speaker in training and test data. Measured across 100 repeated splits, subject-aware splitting drops AUC from 0.93 to 0.80 and specificity from 0.44 to 0.34. Roughly half the apparent skill was speaker recognition. |
Team contribution: coronary risk prediction API, team of seven. My part was feature engineering and exploratory analysis; the two derived features in the deployed model, pulse pressure and smoking intensity, are mine. The API and deployment were built by teammates.
Also on this profile: survival analysis (Kaplan Meier, Cox), applied regression and missing data imputation in R, supervised classification in scikit-learn.
- Languages: R, Python, SQL
- Statistics: multivariable regression, survival analysis, odds ratio and hazard ratio interpretation, missing data and multiple imputation, model validation
- Machine learning: scikit-learn, tree based models, SHAP
- Health data: MIMIC-III (credentialed PhysioNet access), cohort definition, EHR text mining, research data governance under a Data Use Agreement
- Reporting: Power BI, Quarto, R Markdown
- MSc Health Data Science, Universitat Rovira i Virgili, 2025 to 2027
- Marketing Manager, B&F Oberwerth. I built and run the partner and affiliate programme, evaluating 100+ collaborators on performance data.
An internship or placement in 2027, arranged as a convenio de prácticas through URV. Based in Europe.
Roles I am targeting:
- Health data and analytics: data scientist or health data analyst, real world data and HEOR, analytics and insights
- Commercial analytics and business intelligence: BI, commercial analytics, product marketing
- Digital health and AI: associate product manager, AI adoption and implementation, data science delivery
- Email: albers.johanna@outlook.com
- LinkedIn: johanna-albers
- Languages: German, English, Spanish

