Performance of a Large Language Model in BI-RADS Classification of Ultrasound Based Breast Lesions
Pillai, K.; Nausheen, F.
Show abstract
AimsGiven the advent of large language models (LLMs), the number of potential applications using artificial intelligence technologies in radiology has rapidly increased. Recently, several studies have evaluated the accuracy and quality of LLMs to characterize CT and MRI scans. Yet, to our knowledge, there have been few studies that have reported the utility of these models in generating BI-RADS assessment categories. MethodsA breast ultrasound dataset including 256 images from 256 patients manually interpreted and labeled by radiologists according to BI-RADS features and lexicon was used for evaluating Gemini 2.0 Flash. We prompted the model to assess images in individual context windows and tested it with two variations of the original prompt (n = 3). Statistical analyses were then performed comparing the abilities of the model to the ground truth. The receiver operating characteristic-area under the curve (ROC-AUC) analysis was then calculated for each classification type from individual replicates. ResultsWe found that the overall accuracy of Gemini 2.0 was 19.01% in predicting the BI-RADS classification of the breast lesions, and those of each category did not significantly differ from one another. From the ROC-AUC analysis, all category scores ranged from 0.5-0.6, and found that the model performed slightly better at categorizing benign lesions (1-4a), while those of greater probability of malignancy were akin to random chance (4b-5). Furthermore, we found that among incorrect predictions, the model was generally within 1-2 categories away from the true classification, demonstrating a low precision unreliable for realistic clinical usage. ConclusionsThis work highlights the current limitations of artificial intelligence models in classifying clinical images, and further development is required in these technologies before translation into the clinical setting. To our knowledge, this is the first study to report the capabilities of LLMs in performing BI-RADS classification of breast lesions with replicates.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Model uncertainty estimates for deep learning mammographic density prediction using ordinal and classification approaches 94%
- Breast density prediction from low and standard dose mammograms using deep learning: effect of image resolution and model training approach on prediction quality 92%
- Mammographic density assessed using deep learning in women at high risk of developing breast cancer: the effect of weight change on density 91%
Similar papers in this journal
- A Machine Learning Ensemble Based on Radiomics to Predict BI-RADS Category and Reduce the Biopsy Rate of Ultrasound-Detected Suspicious Breast Masses 97%
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 91%
- The NILS study protocol - a retrospective validation study of a preoperative decision-making tool for non-invasive lymph node staging in women with primary breast cancer [ISRCTN14341750] 91%
Similar papers in this journal
- Improved accuracy of breast volume calculation from 3D surface imaging data using statistical shape models 94%
- Classification performance bias between training and test sets in a limited mammography dataset 94%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 93%
Similar papers in this journal
- Breast invasive ductal carcinoma classification on whole slide images with weakly-supervised and transfer learning 95%
- From Variability to Standardization: The Impact of Breast Density on Background Parenchymal Enhancement in Contrast-Enhanced Mammography and the Need for a Structured Reporting System 95%
- piNET: An Automated Proliferation Index Calculator Framework for Ki67 Breast Cancer Images 94%
Similar papers in this journal
- Automated and Manual Quantification of Tumour Cellularity in Digital Slides for Tumour Burden Assessment 94%
- Automated detection of the HER2 gene amplification status in Fluorescence in situ hybridization images for the diagnostics of cancer tissues 92%
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.