Use of Federated Learning for validating and updating privacy-preserving decentralized multi-study prognostic models in Traumatic Brain Injury
Torres-Espin, A.; Wong, J. C.; Hinson, H. E.; Kuipers, T. B.; Hoekstra, B. P. T.; Jain, S.; Sun, X.; Yue, J. K.; Pisi?, D.; Mikolic, A.; Lingsma, H. F.; Markowitz, A. J.; Ferguson, A. R.; Menon, D. K.; Maas, A. I. R.; Steyerberg, E. W.; Manley, G. T.; Belton, P. J.
Show abstract
Developing modern clinical prediction models (CPMs) and advanced analytics requires large datasets, often necessitating data from different studies. Privacy regulations may hinder data sharing, especially across countries. Decentralized federated data infrastructures, where data remain in their original location and analyses are run only in a shared, secure environment, may address these challenges. We implemented a privacy-preserving federated learning (FL) infrastructure and evaluated and updated the IMPACT prognostic models for traumatic brain injury (TBI) using 2 studies. A multi-continental federated infrastructure was established between 2 large-scale studies (TRACK-TBI from the United States and CENTER-TBI from Europe and Israel). Three IMPACT prognostic models for post-TBI 6-month mortality and unfavorable outcomes were evaluated, followed by model updates through 2 FL approaches trained across the TRACK-TBI and CENTER-TBI studies. Internal validation, external cross-validation, and sub-study validations were performed. CPMs were evaluated for discrimination and calibration. The federated cohort included 1616 participants (TRACK-TBI: n=441, CENTER-TBI: n=1175). Both FL performed well, with comparable coefficient estimates, AUCs (area under the receiver operating characteristics curve) between 0.77-0.88, and calibrated probabilities. Compared to the original IMPACT and single-study models, both federated models presented similar discrimination (AUC), were well-calibrated, were more efficient (higher precision), and reduced the impact of missing data in model estimation. FL is feasible for privacy-preserving development and evaluation of CPMs, and can enable validation and updating across large, virtually analyzed datasets while overcoming regulatory constraints on data combination. Federated infrastructures can facilitate global collaboration to advance data-hungry analytical methods, such as artificial intelligence.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Clinical Neuroimaging Platform for Rapid, Automated Lesion Detection and Personalized Post-Stroke Outcome Prediction 92%
- Enhancing Privacy-Preserving Deployable Large Language Models for Perioperative Complication Detection: A Targeted Strategy with LoRA Fine-tuning 92%
- Zero Shot Health Trajectory Prediction Using Transformer 91%
Similar papers in this journal
Similar papers in this journal
- Leveraging Temporal Learning with Dynamic Range (TLDR) for Enhanced Prediction of Outcomes in Recurrent Exposure and Treatment Settings in Electronic Health Records 92%
- Early risk assessment for COVID-19 patients from emergency department data using machine learning 92%
- Synthetic data for privacy-preserving clinical risk prediction 91%
Similar papers in this journal
- Advancing data science in drug development through an innovative computational framework for data sharing and statistical analysis 93%
- Scalable information extraction from free text electronic health records using large language models 88%
- Assessing the transportability of clinical prediction models for cognitive impairment using causal models 88%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.