An end to end analysis of inpatient hospital administrative data in R, written as preparation for a Research Data Scientist technical assessment. The pipeline runs from raw relational tables through cohort construction, data quality auditing, feature engineering and logistic regression to model diagnostics and reporting.
The data is fully synthetic. See Data provenance below before reading any result as a finding.
gemini_template.Rmd Skeleton version with placeholder column names, written to be
filled in against an unfamiliar schema under time pressure.
gemini_reference_answer.Rmd Complete worked analysis, 787 lines.
ipadmdad.csv 5,000 inpatient admissions (encounter level).
ipdiagnosis.csv 18,888 diagnosis rows, ICD-10-CA coded.
lab.csv 64,890 laboratory results, OMOP-mapped test identifiers.
Open gemini_reference_answer.Rmd in RStudio and knit. All paths are relative to this folder.
Three tables linked by genc_id, the encounter identifier.
ipadmdad.csv
| Column | Description |
|---|---|
genc_id |
Encounter identifier, links all three tables |
patient_id_hashed |
De-identified patient identifier |
hospital_num |
Site identifier |
admission_date_time, discharge_date_time |
Admission and discharge timestamps |
age, gender |
Patient demographics |
discharge_disposition |
CIHI discharge disposition code |
number_of_alc_days, alc_service_transfer_flag |
Alternate level of care fields |
ipdiagnosis.csv: genc_id, hospital_num, diagnosis_code (ICD-10-CA), diagnosis_type
lab.csv: genc_id, test_type_mapped_omop, result_value, result_unit, collection_date_time
1. Cohort construction. Four sequential inclusion criteria with an attrition table reporting N and exclusions at each step: adults aged 18 and over, non-missing discharge disposition, length of stay of at least 4 hours, and admission within a three year Ontario fiscal window (April 2019 to March 2022).
2. Data quality and sanity checks. Plausible range checks on continuous variables, age distribution by hospital, a missingness heatmap across sites, and crude mortality rate by hospital as a face validity check.
3. Feature engineering. In-hospital mortality derived from CIHI discharge disposition codes 7, 72 and 73. Length of stay computed from admission and discharge timestamps. Comorbidity flags built from ICD-10-CA prefixes (E10 and E11 for diabetes, N17 and N18 for renal, C for cancer). First-24-hour lab values extracted as one value per test per encounter, pivoted wide, with lab columns identified dynamically and assessed for missingness and plausible range.
4. Descriptive analysis. Table 1 stratified by outcome, monthly mortality trend, age and mortality relationship checked for non-linearity using 5-year bins, length of stay distribution by outcome, lab correlation heatmap, comorbidity co-occurrence.
5. Modelling. Logistic regression for in-hospital mortality with an odds ratio table, variance inflation factors (using GVIF^(1/(2·Df)) for factor terms), ROC with confidence intervals on the AUC, a decile-based calibration plot, and a forest plot. A lab-augmented model is compared to the base model on a common complete-case subset using a likelihood ratio test and a DeLong test for the difference in AUC.
6. Summary and limitations. Written limitations section, with restricted cubic splines for age and multiple imputation for missing labs identified as next steps.
The three CSV files are synthetic data generated by gemSim, the GEMINI team's own synthetic data package.
- Source: https://github.com/GEMINI-Medicine/gemSim
- Licence: MIT. Copyright (c) 2026 GEMINI team, Unity Health Toronto. Full notice in
THIRD_PARTY_NOTICES.mdin this folder. - gemSim does not sample from or reproduce real GEMINI data. It generates fully synthetic records that approximate the database schema and high-level distributional characteristics of the real holdings.
Read this before interpreting any number in the analysis. gemSim states that clinical outcomes, predictors and covariates are simulated independently of one another, and therefore should not be interpreted as having any meaningful associations. Every odds ratio, p-value, AUC and calibration curve in gemini_reference_answer.Rmd is a demonstration that the code runs and the diagnostics fire correctly. None of them is a finding. The artifact here is the pipeline and the reasoning, not the results.
Written as preparation for the GEMINI Research Data Scientist technical assessment, a timed applied analysis task on hospital EHR data. The structure follows GEMINI's relational data architecture and CIHI coding conventions.
This is preparation material and a demonstration pipeline. It is not a study, a publication, or the assessment submission itself.