ABSTRACT
Background Temporal distribution shift negatively impacts the performance of clinical prediction models over time. Pretraining foundation models using self-supervised learning on electronic health records (EHR) may be effective in acquiring informative global patterns that can improve the robustness of task-specific models.
Objective To evaluate the utility of EHR foundation models in improving the in-distribution (ID) and out-of-distribution (OOD) performance of clinical prediction models.
Methods The cohort consisted of adult inpatients admitted between 2009-2021. Gated recurrent unit (GRU)- and transformer (TRANS)-based foundation models were pretrained on EHR of patients admitted between 2009-2012 and were subsequently used to construct patient representations (CLMBR). These representations were used to learn logistic regression models (CLMBRGRU and CLMBRTRANS) to predict hospital mortality, long length of stay, 30-day readmission, and ICU admission. We compared CLMBRGRU and CLMBRTRANS with baseline logistic regression models learned on count-based representations (count-LR) and end-to-end (ETE) GRU and transformer models in ID (2009-2012) and OOD (2013-2021) year groups. Performance was measured using area-under-the-receiver-operating-characteristic curve, area- under-the-precision-recall curve, and absolute calibration error.
Results Models trained on CLMBR generally showed better discrimination relative to count-LR in both ID and OOD year groups. In addition, they often matched or were better than their ETE counterparts. Finally, foundation models’ performance in the self-supervised learning task tracked closely with the ID and OOD performance of the downstream models.
Conclusions These results suggest that pretraining foundation models on electronic health records is a useful approach for developing clinical prediction models that perform well in the presence of temporal distribution shift.
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
No external funding was received for the study.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
This study used de-identified data and so the requirement for Institutional Review Board approval and participant informed consent were waived by Stanford Medical Center.
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.
Yes
Footnotes
↵* co-senior authors
Data Availability
The Stanford Medicine Research Data Repository is not made publicly available. The code for all analyses is open-source and available at https://github.com/som-shahlab/temp_ds_shift_robustness