Detection of Patient Metadata in Published Articles for Genomic Epidemiology Using Machine Learning and Large Language Models
Klein, A. Z.; Weissenbacher, D.; O'Connor, K.; Elyaderani, A.; Flores Amaro, I.; Onishi, T.; Golder, S.; Spiegel, K.; Scotch, M.; Gonzalez-Hernandez, G.
Show abstract
ObjectivePatient metadata exist in published articles, but are often dis-connected from genome sequences in databases, limiting their utility for genomic epidemiology. The objective of this study was to develop and evaluate natural language processing methods to facilitate the large-scale detection of patient metadata associated with reports of genome sequencing in published articles, drawing on the case of SARS-CoV-2. MethodsWe applied filters to select a sample of 245 PubMed articles (50,918 sentences) in LitCovid for manual annotation of sentences that reported generating SARS-CoV-2 sequences. We trained, deployed, and validated a BERT-based classifier, and selected a sample of 150 predicted articles (22,147 sentences) for manual annotation of sentences that reported patient metadata associated with the sequences. In addition to training BERT-based classifiers, we experimented with a generative AI approach, prompting the Llama-3-70B LLM using zero-shot, role-based, few-shot, chain-of-thought, and reasoning-eliciting prompting. ResultsBERT-based models that were pre-trained on corpora in biomedical or, more specifically, COVID-19 domains outperformed those that were pre-trained on corpora in general domains for detecting reports of patient metadata associated with SARS-CoV-2 sequences, achieving the best performance with a classifier based on a BiomedBERT-Large-Abstract model (F1-score = 0.776). While the best performance of our generative AI approach was achieved using role-based, few-shot, and chain-of-thought prompting (F1-score = 0.558), it was nonetheless outperformed by all of our machine learning-based classifiers. ConclusionOur methods were applied to more than 350,000 published articles and can be used to advance the utility and efficiency of genomic epidemiology for public health responses to virus outbreaks.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 94%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 93%
Similar papers in this journal
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 95%
- Machine Learning Maps Research Needs in COVID-19 Literature 94%
- Structuring clinical text with AI: old vs. new natural language processing techniques evaluated on eight common cardiovascular diseases 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.