A hybrid approach to scalable real-world data curation by machine learning and human experts
Waskom, M.; Tan, K.; Wiberg, H.; Cohen, A.; Wittmershaus, B.; Shapiro, W.
Show abstract
ObjectiveMachine learning has the potential to increase the scale of real-world data curated from electronic health records, but maintaining a high standard of data quality is important to avoid biasing downstream analyses. To increase scale without compromising quality, we propose a hybrid data curation methodology that employs both manual abstraction by clinical experts and automated extraction by machine learning models. Materials and MethodsOur methodology makes the determination about when to employ manual abstraction using a confidence score associated with each model output. We describe a process for selecting confidence thresholds based on simulations validated against a reference-standard labeled dataset. To establish the fitness of our methodology for retrospective research, we apply it to a multi-variable cohort selection task on a large real-world oncology database. ResultsOnly small amounts of manual abstraction are required for hybrid curation to achieve expert-level error rates. In fact, the hybrid methodology can even reduce error rates relative to manual abstraction in some cases. We further demonstrate that demographic characteristics of a research cohort defined using hybrid variables are comparable to one curated with conventional methods. DiscussionOur methodology is general and makes few assumptions about the clinical variable or machine learning model. A key requirement is the availability of reference standard labels for calibrating the tradeoff between abstraction effort and data quality. ConclusionIncorporating machine learning into real-world data curation using hybrid methodology holds the promise to scale practicable cohort sizes while maintaining data fitness for research purposes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 94%
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 93%
- Temporal Relationship of Computed and Structured Diagnoses in Electronic Health Record Data 93%
Similar papers in this journal
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 95%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.