Performance assessment of large language models in cancer staging: Comparative analysis of Mistral models
Rouzier, R.; Harter, V.; Rouzier, E.; Ferment, V.; Gruau, S.; Andre, B.; Saumon Sud, C.; Nadin, L.; Corroyer Dulmont, A.; Vigneron, N.
Show abstract
Cancer staging plays a critical role in treatment planning and prognosis but is often embedded in unstructured clinical narratives. To automate the extraction and structuring of staging data, large language models (LLMs) have emerged as a promising approach. However, their performance in real-world oncology settings has yet to be systematically evaluated. Herein, we analysed 1000 oncological summaries from patients receiving treatment for breast cancer between 2019 and 2020 at the Francois Baclesse Comprehensive Cancer Centre, France. Five Mistral artificial intelligence-based LLMs were evaluated (i.e. Small, Medium, Large, Magistral and Mistral:latest) for their ability to derive the cancer stage and identify staging elements. Larger models outperformed their smaller counterparts in staging accuracy and reproducibility (kappa > 0.95 for Mistral Large and Medium). Mistral Large achieved the highest accuracy in deriving the cancer stage (93.0%), surpassing the original clinical documentation in several cases. The LLMs consistently performed better in deriving the cancer stage when working through tumour size, nodal status and metastatic components compared to when they were directly requested stage data. The top-performing models had a test-retest reliability exceeding 97%, while smaller models and locally deployed versions lacked sufficient robustness, particularly in handling unit conversions and complex staging rules. The structured, stepwise use of LLMs that emulates clinician reasoning offers a more efficient, transparent and reproducible approach to cancer staging, and the study findings support LLM integration into digital oncology workflows.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 94%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
- A CDE-based data structure for radiotherapeutic decision-making in breast cancer 93%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 91%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
Similar papers in this journal
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 96%
- Use of natural language understanding to facilitate surgical de-escalation of axillary staging in patients with breast cancer 92%
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 91%
Similar papers in this journal
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 95%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- Stability of feature selection utilizing Graph Convolutional Neural Network and Layer-wise Relevance Propagation 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.