Evaluating Large Language Models in Interpreting Cervical Cytology
Geetha, S. D.
Show abstract
BackgroundLarge language models (LLMs) have shown promise in medical imaging, but their utility in cytology remains underexplored. This study evaluates GPT-5 and Gemini 2.5 Pro for Pap smear interpretation. MethodsDigital cervical Pap smear images of 100 cases were obtained from the Hologic Education Site, with Hologic diagnoses considered the gold standard. Representative images were uploaded into GPT-5 and Gemini 2.5 Pro and prompted to provide a diagnosis based on the Third Edition of the Bethesda System for Reporting Cervical Cytopathology. Cases with infectious organisms were assessed using additional images. Concordance was evaluated at exact diagnosis and clinical management groupings, wherein diagnoses with similar management implications were grouped. Sensitivity and specificity for abnormal cytology were also calculated. ResultsConcordance of both LLMs for exact diagnostic matches were comparable (GPT-5: 47%, Gemini: 48%) and increased to 66% for clinical management grouping. GPT-5 performed best for low-grade squamous intraepithelial lesions (75%), whereas Gemini 2.5 Pro showed the highest concordance in the high-grade squamous intraepithelial lesion (HSIL) category (82%), although this was largely attributable to its strong tendency to overcall cases as HSIL. Sensitivity for detecting abnormal cytology was 74% for GPT-5 and 84% for Gemini, with specificity of 74% and 71%, respectively. GPT-5 better identified glandular lesions, while Gemini detected organisms more accurately (71% vs. 20%). ConclusionsCurrent LLMs demonstrate moderate ability to identify cytologic abnormalities but are not yet reliable for independent Pap smear interpretation. Targeted fine-tuning, prompt optimization, and cytology-specific training could enhance their utility as adjunctive tools in cytology workflows.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial Intelligence for Advance Requesting of Immunohistochemistry in Diagnostically Uncertain Prostate Biopsies 94%
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 93%
- Clinical-Grade Validation of an Autofluorescence Virtual Staining System with Human Experts and a Deep Learning System for Prostate Cancer 93%
Similar papers in this journal
- Using an Anomaly Detection Approach for the Segmentation of Colorectal Cancer Tumors in Whole Slide Images 93%
- Independent assessment of a deep learning system for lymph node metastasis detection on the Augmented Reality Microscope 93%
- Creating virtual H&E images using samples imaged on a commercial CODEX platform 93%
Similar papers in this journal
- Lymphoid Enhancer-Binding Factor 1 (LEF1) immunostaining as a surrogate of β-catenin ( CTNNB1) mutations 90%
- Outcome of radial scar/complex sclerosing lesion associated with epithelial proliferations with atypia diagnosed on breast core biopsy - Results from a multicentric UK based study 89%
- Increasing both specificity and sensitivity of SARS-CoV-2 antibody tests by using an adaptive orthogonal testing approach 87%
Similar papers in this journal
- A Machine Learning Ensemble Based on Radiomics to Predict BI-RADS Category and Reduce the Biopsy Rate of Ultrasound-Detected Suspicious Breast Masses 91%
- Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images using transfer learning 90%
- Demarcation line determination for diagnosis of gastric cancer disease range using unsupervised machine learning in magnifying narrow-band imaging 89%
Similar papers in this journal
- Weakly supervised learning for multi-organ adenocarcinoma classification in whole slide images 92%
- Classification performance bias between training and test sets in a limited mammography dataset 91%
- Pixelwise H-score: a novel digital image analysis based-metric to quantify membrane biomarker expression from immunohistochemistry images 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.