Prompts to Table: Specification and Iterative Refinement for Clinical Information Extraction with Large Language Models
Hein, D.; Christie, A.; Holcomb, M.; Xie, B.; Jain, A.; Vento, J.; Rakheja, N.; Hamza Shakur, A.; Christley, S.; Cowell, L. G.; Brugarolas, J.; Jamieson, A. R.; Kapur, P.
Show abstract
Extracting structured data from free-text medical records at scale is laborious, and traditional approaches struggle in complex clinical domains. We present a novel, end-to-end pipeline leveraging large language models (LLMs) for highly accurate information extraction and normalization from unstructured pathology reports, focusing initially on kidney tumors. Our innovation combines flexible prompt templates, the direct production of analysis-ready tabular data, and a rigorous, human-in-the-loop iterative refinement process guided by a comprehensive error ontology. Applying the finalized pipeline to 2,297 kidney tumor reports with pre-existing templated data available for validation yielded a macro-averaged F1 of 0.99 for six kidney tumor subtypes and 0.97 for detecting kidney metastasis. We further demonstrate flexibility with multiple LLM backbones and adaptability to new domains utilizing publicly available breast and prostate cancer reports. Beyond performance metrics or pipeline specifics, we emphasize the critical importance of task definition, interdisciplinary collaboration, and complexity management in LLM-based clinical workflows.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 93%
- Artificial Intelligence for Advance Requesting of Immunohistochemistry in Diagnostically Uncertain Prostate Biopsies 93%
- Genomic Characterization of Lung Cancer in Never-Smokers Using Deep Learning 93%
Similar papers in this journal
- Development of an Interactive Web Dashboard to Facilitate the Reexamination of Pathology Reports for Instances of Underbilling of CPT Codes 95%
- A tool for federated training of segmentation models on whole slide images. 93%
- Deep learning accurately quantifies plasma cell percentages on CD138-stained bone marrow samples 92%
Similar papers in this journal
- Natural language inference for clinical registry curation 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 91%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 93%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 91%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.