Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

README.md

Cohort Analysis Pipeline on Synthetic GEMINI Hospital Data

An end to end analysis of inpatient hospital administrative data in R, written as preparation for a Research Data Scientist technical assessment. The pipeline runs from raw relational tables through cohort construction, data quality auditing, feature engineering and logistic regression to model diagnostics and reporting.

The data is fully synthetic. See Data provenance below before reading any result as a finding.

Files

gemini_template.Rmd        Skeleton version with placeholder column names, written to be
                           filled in against an unfamiliar schema under time pressure.
gemini_reference_answer.Rmd Complete worked analysis, 787 lines.
ipadmdad.csv               5,000 inpatient admissions (encounter level).
ipdiagnosis.csv            18,888 diagnosis rows, ICD-10-CA coded.
lab.csv                    64,890 laboratory results, OMOP-mapped test identifiers.

Open gemini_reference_answer.Rmd in RStudio and knit. All paths are relative to this folder.

Data schema

Three tables linked by genc_id, the encounter identifier.

ipadmdad.csv

Column Description
genc_id Encounter identifier, links all three tables
patient_id_hashed De-identified patient identifier
hospital_num Site identifier
admission_date_time, discharge_date_time Admission and discharge timestamps
age, gender Patient demographics
discharge_disposition CIHI discharge disposition code
number_of_alc_days, alc_service_transfer_flag Alternate level of care fields

ipdiagnosis.csv: genc_id, hospital_num, diagnosis_code (ICD-10-CA), diagnosis_type

lab.csv: genc_id, test_type_mapped_omop, result_value, result_unit, collection_date_time

What the pipeline does

1. Cohort construction. Four sequential inclusion criteria with an attrition table reporting N and exclusions at each step: adults aged 18 and over, non-missing discharge disposition, length of stay of at least 4 hours, and admission within a three year Ontario fiscal window (April 2019 to March 2022).

2. Data quality and sanity checks. Plausible range checks on continuous variables, age distribution by hospital, a missingness heatmap across sites, and crude mortality rate by hospital as a face validity check.

3. Feature engineering. In-hospital mortality derived from CIHI discharge disposition codes 7, 72 and 73. Length of stay computed from admission and discharge timestamps. Comorbidity flags built from ICD-10-CA prefixes (E10 and E11 for diabetes, N17 and N18 for renal, C for cancer). First-24-hour lab values extracted as one value per test per encounter, pivoted wide, with lab columns identified dynamically and assessed for missingness and plausible range.

4. Descriptive analysis. Table 1 stratified by outcome, monthly mortality trend, age and mortality relationship checked for non-linearity using 5-year bins, length of stay distribution by outcome, lab correlation heatmap, comorbidity co-occurrence.

5. Modelling. Logistic regression for in-hospital mortality with an odds ratio table, variance inflation factors (using GVIF^(1/(2·Df)) for factor terms), ROC with confidence intervals on the AUC, a decile-based calibration plot, and a forest plot. A lab-augmented model is compared to the base model on a common complete-case subset using a likelihood ratio test and a DeLong test for the difference in AUC.

6. Summary and limitations. Written limitations section, with restricted cubic splines for age and multiple imputation for missing labs identified as next steps.

Data provenance

The three CSV files are synthetic data generated by gemSim, the GEMINI team's own synthetic data package.

  • Source: https://github.com/GEMINI-Medicine/gemSim
  • Licence: MIT. Copyright (c) 2026 GEMINI team, Unity Health Toronto. Full notice in THIRD_PARTY_NOTICES.md in this folder.
  • gemSim does not sample from or reproduce real GEMINI data. It generates fully synthetic records that approximate the database schema and high-level distributional characteristics of the real holdings.

Read this before interpreting any number in the analysis. gemSim states that clinical outcomes, predictors and covariates are simulated independently of one another, and therefore should not be interpreted as having any meaningful associations. Every odds ratio, p-value, AUC and calibration curve in gemini_reference_answer.Rmd is a demonstration that the code runs and the diagnostics fire correctly. None of them is a finding. The artifact here is the pipeline and the reasoning, not the results.

Context

Written as preparation for the GEMINI Research Data Scientist technical assessment, a timed applied analysis task on hospital EHR data. The structure follows GEMINI's relational data architecture and CIHI coding conventions.

This is preparation material and a demonstration pipeline. It is not a study, a publication, or the assessment submission itself.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages