Back

Automated ICD-O-3 Coding of Real-World Pathology Reports Using Self-Hosted Large Language Models

Arzideh, K.; Hosch, R.; Turki, A.; Erilymaz, B.; Bahn, M.; Schaefer, H.; Idrissi-Yaghir, A.; Khattab, S.; Dada, A.; Baba, H. A.; Schadendorf, D.; Schuler, M.; Kleesiek, J.; Hartmann, S.; Nensa, F.; Keyl, J.

2025-07-23 pathology
10.1101/2025.07.23.25332040 medRxiv
Show abstract

While large language models (LLMs) have shown promise in medical text processing, their real-world application in self-hosted clinical settings remains underexplored. Here, we evaluated five state-of-the-art self-hosted LLMs for automated assignment of International Classification of Diseases for Oncology (ICD-O-3) codes using 21,364 real-world pathology reports from a large German hospital. For exact topography code prediction, Qwen3-235B-A22B achieved the highest performance (micro-average F1: 71.6%), while Llama-3.3-70B-Instruct performed best score at predicting the first three characters (micro-average F1: 84.6%). For morphology codes, DeepSeek-R1-Distill-Llama-70B outperformed other models (exact micro-average F1: 34.7%; first three characters micro-average F1: 77.8%). Large disparities between micro- and macro-average F1-scores indicated poor generalization to rare conditions. Although LLMs demonstrate promising capabilities as support systems for expert-guided pathology coding, their performance is not yet sufficient for fully automated, unsupervised use in routine clinical workflows.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.