Accuracy of Foundation AI Models for Hepatic Macrovesicular Steatosis Quantification in Frozen Sections
Koga, S.; Guda, A.; Wang, Y.; Sahni, A.; Wu, J.; Rosen, A.; Nield, J.; Nandish, N.; Patel, K.; Goldman, H.; Rajapakse, C.; Walle, S.; Kristen, S.; Tondon, R.; Alipour, Z.
Show abstract
IntroductionAccurate intraoperative assessment of macrovesicular steatosis in donor liver biopsies is critical for transplantation decisions but is often limited by inter-observer variability and freezing artifacts that can obscure histological details. Artificial intelligence (AI) offers a potential solution for standardized and reproducible evaluation. To evaluate the diagnostic performance of two self-supervised learning (SSL)-based foundation models, Prov-GigaPath and UNI, for classifying macrovesicular steatosis in frozen liver biopsy sections, compared with assessments by surgical pathologists. MethodsWe retrospectively analyzed 131 frozen liver biopsy specimens from 68 donors collected between November 2022 and September 2024. Slides were digitized into whole-slide images, tiled into patches, and used to extract embeddings with Prov-GigaPath and UNI; slide-level classifiers were then trained and tested. Intraoperative diagnoses by on-call surgical pathologists were compared with ground truth determined from independent reviews of permanent sections by two liver pathologists. Accuracy was evaluated for both five-category classification and a clinically significant binary threshold (<30% vs. [≥]30%). ResultsFor binary classification, Prov-GigaPath achieved 96.4% accuracy, UNI 85.7%, and surgical pathologists 84.0% (P = .22). In five-category classification, accuracies were lower: Prov-GigaPath 57.1%, UNI 50.0%, and pathologists 58.7% (P = .70). Misclassification primarily occurred in intermediate categories (5%-<30% steatosis). ConclusionsSSL-based foundation models performed comparably to surgical pathologists in classifying macrovesicular steatosis, at the clinically relevant <30% vs. [≥]30% threshold. These findings support the potential role of AI in standardizing intraoperative evaluation of donor liver biopsies; however, the small sample size limits generalizability and requires validation in larger, balanced cohorts.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MIXTURE of human expertise and deep learning—Developing an explainable model for predicting pathological diagnosis and survival in patients with interstitial lung disease 91%
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 90%
- Artificial Intelligence for Advance Requesting of Immunohistochemistry in Diagnostically Uncertain Prostate Biopsies 89%
Similar papers in this journal
Similar papers in this journal
- Predicting Short-Term Mortality in Severe Cirrhosis: An Interpretable Machine Learning Model Integrating Routine Clinical Indicators 92%
- LIMPACAT : Multi-Omics Attention Transformer for Immune Prediction in Liver Cancer Using Whole-Slide Imaging 91%
- Detection of infiltrating fibroblasts by single-cell transcriptomics in human kidney allografts 91%
Similar papers in this journal
- Enhancing Liver Fibrosis Measurement: Deep Learning and Uncertainty Analysis Across Multi-Centre Cohorts 92%
- Using an Anomaly Detection Approach for the Segmentation of Colorectal Cancer Tumors in Whole Slide Images 90%
- Deep learning accurately quantifies plasma cell percentages on CD138-stained bone marrow samples 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.