LLM-based data extraction for a large cancer registry, the Ontario Hereditary Cancer Research Network
Melani De La Hoz, A. F.; Weile, J.; Hemlani, P.; Tuzlali, E.; Chan, B.; Chun, K.; Feilotter, H.; Grafodatskaya, D.; Hughes, L. K.; Kim, R. H.; Lerner-Ellis, J.; Ridd, S.; Schenkel, L.; Smith, A.; Stein, L.; Vaags, A.; Wang, H.; Haibe-Kains, B.; Courtot, M.
Show abstract
ImportanceManual data extraction from genomic lab reports for on-line registries and databases is time-consuming for human resources such as clinical research coordinators. Automated tools, especially LLMs, can address these issues. Efficient and accurate data processing is crucial for building a reliable database. ObjectiveTo streamline the data extraction and curation process for genetic testing lab reports using an LLM-based approach. DesignNine sample molecular lab reports were selected for manual data extraction by two expert curators. The process was timed, and the results served as gold-standard for validating automated extraction. Eighteen fields from the OHRCNs data model were selected as extraction targets. SettingThe study was conducted within OHCRN, which unifies research, genomic, and clinical patient data from clinics and laboratories across Ontario, Canada. ParticipantsNine laboratories agreed to share sample molecular lab reports and two clinical research coordinators affiliated with OHCRN participated as data curators. ExposureLLM-based Extraction of Information (LEI), an automated data extraction pipeline, was developed using regular expressions, Trie search, and LLMs to extract data from molecular lab reports and structure it for inclusion into OHCRNs database. Main Outcomes and MeasuresLEI was evaluated by measuring the F1-score on the extraction task of 18 entity types. These measures were compared against 15 extraction tools in the biomedical domain. Extraction time was also measured and compared against manual extraction times. ResultsLEI demonstrated quality on par with and surpassing other existing LLM-based extraction methods. Reference tools showed F1-scores around 70%, while LEI achieved an average score of 87.4%. LEI reduced extraction time by approximately 2-fold, with an average time of 7.59 minutes per report including results review by curators, compared to 14.88 minutes per report for manual extraction. Conclusions and RelevanceLEI facilitates standardized, accurate, and efficient healthcare data extraction from unstructured texts, significantly improving the current OHCRN workflow. By automating the extraction process, LEI allows expert curators to focus on validating results rather than performing manual data entry. LEIs simple interface enables researchers to easily guide extraction tasks and supports adaptability across diverse biomedical scenarios. Future improvements in accuracy may be achieved through fine-tuning techniques and ongoing advancements in LLM technologies.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 91%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 90%
Similar papers in this journal
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 93%
- Systematic identification of rare disease patients in electronic health records enables evaluation of clinical outcomes 93%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 92%
Similar papers in this journal
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 94%
- Large-Scale Deep Learning for Metastasis Detection in Pathology Reports 93%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.