Development and Validation of Natural Language Processing Algorithms in the ENACT National Electronic Health Record Research Network
Wang, Y.; Hilsman, J.; Li, C.; Morris, M.; Heider, P. M.; Fu, S.; Kwak, M. J.; Wen, A.; Applegate, J. R.; Wang, L.; Bernstam, E. V.; Liu, H.; Chang, J.; Harris, D. R.; Corbeau, A.; Henderson, D.; Osborne, J. D.; Kennedy, R. E.; Garduno-Rapp, N.-E.; Rousseau, J. F.; Yan, C.; Chen, Y.; Patel, M. B.; Murphy, T. J.; Malin, B. A.; Park, C. M.; Fan, J. W.; Sohn, S.; Pagali, S.; Peng, Y.; Pathak, A.; Wu, Y.; Xia, Z.; Loguercio, S.; Reis, S. E.; Visweswaran, S.
Show abstract
Electronic health record (EHR) data are a rich and invaluable source of real-world clinical information, enabling detailed insights into patient populations, treatment outcomes, and healthcare practices. The availability of large volumes of EHR data are critical for advancing translational research and developing innovative technologies such as artificial intelligence. The Evolve to Next-Gen Accrual to Clinical Trials (ENACT) network, established in 2015 with funding from the National Center for Advancing Translational Sciences (NCATS), aims to accelerate translational research by democratizing access to EHR data for all Clinical and Translational Science Awards (CTSA) hub investigators. The present ENACT network provides access to structured EHR data, enabling cohort discovery and translational research across the network. However, a substantial amount of critical information is contained in clinical narratives, and natural language processing (NLP) is required for extracting this information to support research. To address this need, the ENACT NLP Working Group was formed to make NLP-derived clinical information accessible and queryable across the network. This article describes the implementation and deployment of NLP infrastructure across ENACT. First, we describe the formation and goals of the Working Group, the practices and logistics involved in implementation and deployment, and the specific NLP tools and technologies utilized. Then, we describe how we extended the ENACT ontology to standardize and query NLP-derived data, as well as how we conducted multisite evaluations of the NLP algorithms. Finally, we reflect on the experience and lessons learnt, which may be useful for other national data networks that are deploying NLP to unlock the potential of clinical text for research.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Development of a COVID-19 Application Ontology for the ACT Network 96%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 96%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 97%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 95%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 95%
Similar papers in this journal
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 95%
- Extracting social determinants of health from electronic health records: development and comparison of rule-based and large language models-based methods 94%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.