Generating highly accurate pathology reports from gigapixel whole slide images with HistoGPT
Tran, M.; Schmidle, P.; Wagner, S. J.; Koch, V.; Lupperger, V.; Feuchtinger, A.; Boehner, A.; Kaczmarczyk, R.; Biedermann, T.; Eyerich, K.; Braun, S. A.; Peng, T.; Marr, C.
Show abstract
Histopathology is considered the reference standard for diagnosing the presence and nature of many malignancies, including cancer. However, analyzing tissue samples and writing pathology reports is time-consuming, labor-intensive, and non-standardized. To address this problem, we present HistoGPT, the first vision language model that simultaneously generates reports from multiple pathology images. It was trained on more than 15,000 whole slide images from over 6,000 dermatology patients with corresponding pathology reports. The generated reports match the quality of human-written reports, as confirmed by a variety of natural language processing metrics and domain expert evaluations. We show that HistoGPT generalizes to six geographically diverse cohorts and can predict tumor subtypes and tumor thickness in a zero-shot fashion. Our model demonstrates the potential of an AI assistant that supports pathologists in evaluating, reporting, and understanding routine dermatopathology cases.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 95%
- Generative AI Enables Medical Image Segmentation in Ultra Low-Data Regimes 95%
- PHARAOH: A collaborative crowdsourcing platform for PHenotyping And Regional Analysis Of Histology 95%
Similar papers in this journal
Similar papers in this journal
- Assessing large multimodal models for one-shot learning and interpretability in biomedical image classification 96%
- Label-free virtual peritoneal lavage cytology via deep-learning-assisted single-color stimulated Raman scattering microscopy 93%
- Looming detection in complex dynamic visual scenes by interneuronal coordination of motion and feature pathways 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.