Back

Cumulus: A federated EHR-based learning system powered by FHIR and AI

McMurry, A. J.; Gottlieb, D. I.; Miller, T. A.; Jones, J. R.; Atreja, A.; Crago, J.; Desai, P. M.; Dixon, B. E.; Garber, M.; Ignatov, V.; Kirchner, L. A.; Payne, P. R.; Saldanha, A. J.; Shankar, P. R.; Solad, Y. V.; Sprouse, E. A.; Terry, M.; Wilcox, A. B.; Mandl, K. D.

2024-02-06 health informatics
10.1101/2024.02.02.24301940 medRxiv
Show abstract

ObjectiveTo address challenges in large-scale electronic health record (EHR) data exchange, we sought to develop, deploy, and test an open source, cloud-hosted app listener that accesses standardized data across the SMART/HL7 Bulk FHIR Access application programming interface (API). MethodsWe advance a model for scalable, federated, data sharing and learning. Cumulus software is designed to address key technology and policy desiderata including local utility, control, and administrative simplicity as well as privacy preservation during robust data sharing, and AI for processing unstructured text. ResultsCumulus relies on containerized, cloud-hosted software, installed within a healthcare organizations security envelope. Cumulus accesses EHR data via the Bulk FHIR interface and streamlines automated processing and sharing. The modular design enables use of the latest AI and natural language processing tools and supports provider autonomy and administrative simplicity. In an initial test, Cumulus was deployed across five healthcare systems each partnered with public health. Cumulus output is patient counts which were aggregated into a table stratifying variables of interest to enable population health studies. All code is available open source. A policy stipulating that only aggregate data leave the institution greatly facilitated data sharing agreements. Discussion and ConclusionCumulus addresses barriers to data sharing based on (1) federally required support for standard APIs (2), increasing use of cloud computing, and (3) advances in AI. There is potential for scalability to support learning across myriad network configurations and use cases.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.