Configuring a federated network of real-world patient health data for multimodal deep learning prediction of health outcomes.
Haudenschild, C.; Vaickus, L.; Levy, J.
Show abstract
Vast quantities of electronic patient medical data are currently being collated and processed in large federated data repositories. For instance, TriNetX, Inc., a global health research network, has access to more than 300 million patients, sourced from healthcare organizations, biopharmaceutical companies, and contract research organizations. As such, pipelines that are able to algorithmically extract huge quantities of patient data from multiple modalities present opportunities to leverage machine learning and deep learning approaches with the possibility of generating actionable insight. In this work, we present a modular, semi-automated end-to-end machine and deep learning pipeline designed to interface with a federated network of structured patient data. This proof-of-concept pipeline is disease-agnostic, scalable, and requires little domain expertise and manual feature engineering in order to quickly produce results for the case of a user-defined binary outcome event. We demonstrate the pipelines efficacy with three different disease workflows, with high discriminatory power achieved in all cases.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 95%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- Explainable deep learning for disease activity prediction in chronic inflammatory joint diseases 93%
Similar papers in this journal
Similar papers in this journal
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 92%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 92%
- Generative AI Mitigates Representation Bias Using Synthetic Health Data 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.