Identifying secondary findings in PET/CT reports in oncological cases: A quantifying study using automated Natural Language Processing
Sekler, J.; Kämpgen, B.; Reinert, C. P.; Daul, A.; Gückel, B.; Dittmann, H.; Pfannenberg, C.; Gatidis, S.
Show abstract
BackgroundBecause of their accuracy, positron emission tomography/computed tomography (PET/CT) examinations are ideally suited for the identification of secondary findings but there are only few quantitative studies on the frequency and number of those. Most radiology reports are freehand written and thus secondary findings are not presented as structured evaluable information and the effort to manually extract them reliably is a challenge. Thus we report on the use of natural language processing (NLP) to identify secondary findings from PET/CT conclusions. Methods4,680 anonymized German PET/CT radiology conclusions of five major primary tumor entities were included in this study. Using a commercially available NLP tool, secondary findings were annotated in an automated approach. The performance of the algorithm in classifying primary diagnoses was evaluated by statistical comparison to the ground truth as recorded in the patient registry. Accuracy of automated classification of secondary findings within the written conclusions was assessed in comparison to a subset of manually evaluated conclusions. ResultsThe NLP method was evaluated twice. First, to detect the previously known principal diagnosis, with an F1 score between 0.65 and 0.95 among 5 different principal diagnoses. Second, affirmed and speculated secondary diagnoses were annotated, and the error rate of false positives and false negatives was evaluated. Overall, rates of false-positive findings (1.0%-5.8%) and misclassification (0%-1.1%) were low compared with the overall rate of annotated diagnoses. Error rates for false-negative annotations ranged from 6.1% to 24%. More often, several secondary findings were not fully captured in a conclusion. This error rate ranged from 6.8% to 45.5%. ConclusionsNLP technology can be used to analyze unstructured medical data efficiently and quickly from radiological conclusions, despite the complexity of human language. In the given use case, secondary findings were reliably found in in PET/CT conclusions from different main diagnoses.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 95%
- Predicting EGFR mutation status in lung adenocarcinoma presenting as ground-glass opacity: utilizing radiomics model in clinical translation 93%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 93%
Similar papers in this journal
- Segmentation stability of human head and neck medical images for radiotherapy applications under de-identification conditions: benchmarking for data sharing and artificial intelligence use-cases 91%
- Predicting Axillary Lymph Node Metastasis in Early Breast Cancer Using Deep Learning on Primary Tumor Biopsy Slides 91%
- Deep-Learning-Based Generation of Synthetic High-Resolution MRI from Low-Resolution MRI for Use in Head and Neck Cancer Adaptive Radiotherapy 91%
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 95%
- Weakly supervised learning for multi-organ adenocarcinoma classification in whole slide images 94%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 93%
Similar papers in this journal
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- Identifying relationships between imaging phenotypes and lung cancer-related mutation status: EGFR and KRAS 94%
- High-Dimensional Multinomial Multiclass Severity Scoring of COVID-19 Pneumonia Using CT Radiomics Features and Machine Learning Algorithms 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.