Agent-Based Large Language Model System for Extracting Structured Data from Breast Cancer Synoptic Reports: A Dual-Validation Study
Hart, S. N.; Bergamaschi, T. S.
Show abstract
ObjectiveTo develop and validate an agent-based Large Language Model (LLM) system for extracting structured data from breast cancer synoptic pathology reports and assess the performance gap between synthetic and real-world validation. Materials and MethodsWe developed a modular AI agent-based framework employing sequential specialized LLMs for parsing pathology reports and extracting structured data. We normalized College of American Pathologists (CAP) cancer protocols into 8 sections, 86 subsections, and 229 discrete fields. Seven leading LLMs (gemini-2.5-pro, llama3.3-70b, phi4-14b, deepseek-r1 14B/70B, gemma3-27b, gemini-2.0-flash-lite) were validated using dual evaluation: synthetic validation (864 controlled test cases) and real-world ground truth (6,651 annotated fields from 90 pathology reports). ResultsSynthetic validation demonstrated strong performance (accuracy: 93.8-99.0%). Real-world evaluation revealed field extraction accuracy ranging from 61.8% to 87.7%, demonstrating a substantial "reality gap" with accuracy drops of 11-32 percentage points. The gemini-2.5-pro model achieved the highest real-world accuracy (87.7%). Model size did not predict performance: the 14B-parameter deepseek-r1 (77.6%) outperformed its 70B-parameter counterpart (70.4%). DiscussionThe substantial performance degradation from synthetic to real-world data underscores the complexity of authentic clinical documentation. Smaller models can achieve competitive or superior accuracy, reducing computational costs. With even the best models missing 12-38% of annotated fields, mandatory human verification is essential for clinical deployment. ConclusionWhile LLM-based extraction systems show promise for pathology data extraction, synthetic validation alone provides false confidence. Rigorous real-world ground truth evaluation with expert annotation is essential before clinical deployment. These systems are best positioned as screening tools with mandatory human oversight rather than autonomous decision-making systems.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- Large-Scale Deep Learning for Metastasis Detection in Pathology Reports 95%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 92%
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 92%
Similar papers in this journal
- From months to minutes: creating Hyperion, a novel data management system expediting data insights for oncology research and patient care 93%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 93%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 93%
Similar papers in this journal
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 96%
- Use of natural language understanding to facilitate surgical de-escalation of axillary staging in patients with breast cancer 93%
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.