Back

European Radiology

Springer Science and Business Media LLC

All preprints, ranked by how well they match European Radiology's content profile, based on 15 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Machine learning-based noise reduction for craniofacial bone segmentation in CT images

Nishimoto, S.; Saito, T.; Ishise, H.; Fujiwara, T.; Kawai, K.; Kakibuchi, M.

2022-06-27 dentistry and oral medicine 10.1101/2022.06.26.22276925 medRxiv
Top 0.1%
41.1%
Show abstract

ObjectiveWhen high X-ray absorption rate materials such as metal prosthetics are in the field of CT scan, noise called metal artifacts might appear. In reconstructing a three-dimensional bone model from X-ray CT images, the metal artifacts remain. Often, the image of the scanning bed also remains. A machine learning-based system to reduce noises in the craniofacial CT images was constructed. MethodsDICOM images of CT archives of patients with head and neck tumors were used. The metal artifacts and beds were removed from the threshold segmented images to obtain the target bony images. U-nets, respectively with the function loss of mean squared error, Dice and Jaccard, were trained by the datasets consisting of 5671 DICOM images and corresponding target images. DICOM images of 2000 validation datasets were given to the trained models and predicted images were obtained. ResultsThe use of mean squared errors presented superiority to Dice or Jaccard loss. The mean prediction error pixels were 14.43, 778.57, and 757.60 respectively per 512 x 512 pixeled image. DiscussionDedicated to the delineation of craniofacial bones, the presented study showed high prediction accuracy. The simplification of the target images may have contributed. The "correctness" of the predictions made by this system may not be guaranteed, but the predictions made in this system were generally satisfactory.

2
A Comparative Study: Diagnostic Performance of ChatGPT 3.5, Google Bard, Microsoft Bing, and Radiologists in Thoracic Radiology Cases

Gunes, Y. C.; Cesur, T.

2024-01-20 radiology and imaging 10.1101/2024.01.18.24301495 medRxiv
Top 0.1%
39.1%
Show abstract

PurposeTo investigate and compare the diagnostic performance of ChatGPT 3.5, Google Bard, Microsoft Bing, and two board-certified radiologists in thoracic radiology cases published by The Society of Thoracic Radiology. Materials and MethodsWe collected 124 "Case of the Month" from the Society of Thoracic Radiology website between March 2012 and December 2023. Medical history and imaging findings were input into ChatGPT 3.5, Google Bard, and Microsoft Bing for diagnosis and differential diagnosis. Two board-certified radiologists provided their diagnoses. Cases were categorized anatomically (parenchyma, airways, mediastinum-pleura-chest wall, and vascular) and further classified as specific or non-specific for radiological diagnosis. Diagnostic accuracy and differential diagnosis scores were analyzed using chi-square, Kruskal-Wallis and Mann-Whitney U tests. ResultsAmong 124 cases, ChatGPT demonstrated the highest diagnostic accuracy (53.2%), outperforming radiologists (52.4% and 41.1%), Bard (33.1%), and Bing (29.8%). Specific cases revealed varying diagnostic accuracies, with Radiologist I achieving (65.6%), surpassing ChatGPT (63.5%), Radiologist II (52.0%), Bard (39.5%), and Bing (35.4%). ChatGPT 3.5 and Bing had higher differential scores in specific cases (P<0.05), whereas Bard did not (P=0.114). All three had a higher diagnostic accuracy in specific cases (P<0.05). No differences were found in the diagnostic accuracy or differential diagnosis scores of the four anatomical location (P>0.05). ConclusionChatGPT 3.5 demonstrated higher diagnostic accuracy than Bing, Bard and radiologists in text-based thoracic radiology cases. Large language models hold great promise in this field under proper medical supervision.

3
Performance and Robustness of Machine Learning-based Radiomic COVID-19 Severity Prediction

Yip, S. S. F.; Klanecek, Z.; Naganawa, S.; Kim, J.; Studen, A.; Rivetti, L.; Jeraj, R.

2020-09-09 radiology and imaging 10.1101/2020.09.07.20189977 medRxiv
Top 0.1%
34.5%
Show abstract

ObjectivesThis study investigated the performance and robustness of radiomics in predicting COVID-19 severity in a large public cohort. MethodsA public dataset of 1110 COVID-19 patients (1 CT/patient) was used. Using CTs and clinical data, each patient was classified into mild, moderate, and severe by two observers: (1) dataset provider and (2) a board-certified radiologist. For each CT, 107 radiomic features were extracted. The dataset was randomly divided into a training (60%) and holdout validation (40%) set. During training, features were selected and combined into a logistic regression model for predicting severe cases from mild and moderate cases. The models were trained and validated on the classifications by both observers. AUC quantified the predictive power of models. To determine model robustness, the trained models was cross-validated on the inter-observers classifications. ResultsA single feature alone was sufficient to predict mild from severe COVID-19 with [Formula] and [Formula] (p<< 0.01). The most predictive features were the distribution of small size-zones (GLSZM-SmallAreaEmphasis) for providers classification and linear dependency of neighboring voxels (GLCM-Correlation) for radiologists classification. Cross-validation showed that both [Formula]. In predicting moderate from severe COVID-19, first-order-Median alone had sufficient predictive power of [Formula]. For radiologists classification, the predictive power of the model increased to [Formula] as the number of features grew from 1 to 5. Cross-validation yielded [Formula] and [Formula]. ConclusionsRadiomics significantly predicted different levels of COVID-19 severity. The prediction was moderately sensitive to inter-observer classifications, and thus need to be used with caution. Key pointsO_LIInterpretable radiomic features can predict different levels of COVID-19 severity C_LIO_LIMachine Learning-based radiomic models were moderately sensitive to inter-observer classifications, and thus need to be used with caution C_LI

4
Comparative Analysis of ChatGPT's Diagnostic Performance with Radiologists Using Real-World Radiology Reports of Brain Tumors

Mitsuyama, Y.; Tatekawa, H.; Takita, H.; Sasaki, F.; Tashiro, A.; Satoshi, O.; Walston, S. L.; Miki, Y.; Ueda, D.

2023-10-28 radiology and imaging 10.1101/2023.10.27.23297585 medRxiv
Top 0.1%
32.4%
Show abstract

BackgroundLarge Language Models like Chat Generative Pre-trained Transformer (ChatGPT) have demonstrated potential for differential diagnosis in radiology. Previous studies investigating this potential primarily utilized quizzes from academic journals, which may not accurately represent real-world clinical scenarios. PurposeThis study aimed to assess the diagnostic capabilities of ChatGPT using actual clinical radiology reports of brain tumors and compare its performance with that of neuroradiologists and general radiologists. MethodsWe consecutively collected brain MRI reports from preoperative brain tumor patients at Osaka Metropolitan University Hospital, taken from January to December 2021. ChatGPT and five radiologists were presented with the same findings from the reports and asked to suggest differential and final diagnoses. The pathological diagnosis of the excised tumor served as the ground truth. Chi-square tests and Fishers exact test were used for statistical analysis. ResultsIn a study analyzing 99 radiological reports, ChatGPT achieved a final diagnostic accuracy of 75% (95% CI: 66, 83%), while radiologists accuracy ranged from 64% to 82%. ChatGPTs final diagnostic accuracy using reports from neuroradiologists was higher at 82% (95% CI: 71, 89%), compared to 52% (95% CI: 33, 71%) using those from general radiologists with a p-value of 0.012. In the realm of differential diagnoses, ChatGPTs accuracy was 95% (95% CI: 91, 99%), while radiologists fell between 74% and 88%. Notably, for these differential diagnoses, ChatGPTs accuracy remained consistent whether reports were from neuroradiologists (96%, 95% CI: 89, 99%) or general radiologists (91%, 95% CI: 73, 98%) with a p-value of 0.33. ConclusionChatGPT exhibited good diagnostic capability, comparable to neuroradiologists in differentiating brain tumors from MRI reports. ChatGPT can be a second opinion for neuroradiologists on final diagnoses and a guidance tool for general radiologists and residents, especially for understanding diagnostic cues and handling challenging cases. SummaryThis study evaluated ChatGPTs diagnostic capabilities using real-world clinical MRI reports from brain tumor cases, revealing that its accuracy in interpreting brain tumors from MRI findings is competitive with radiologists. Key resultsO_LIChatGPT demonstrated a diagnostic accuracy rate of 75% for final diagnoses based on preoperative MRI findings from 99 brain tumor cases, competing favorably with five radiologists whose accuracies ranged between 64% and 82%. For differential diagnoses, ChatGPT achieved a remarkable 95% accuracy, outperforming several of the radiologists. C_LIO_LIRadiology reports from neuroradiologists and general radiologists showed varying accuracy when input into ChatGPT. Reports from neuroradiologists resulted in higher diagnostic accuracy for final diagnoses, while there was no difference in accuracy for differential diagnoses between neuroradiologists and general radiologists. C_LI

5
Assessment of the accuracy of lung lesions diagnosis in adolescents with osteosarcoma using artificial intelligence

Uskova, N. G.; Gombolevskiy, V. A.; Chernina, V. Y.; Burenchev, D. V.; Akhaladze, D. G.; Panina, E. V.; Karachunskiy, A. I.; Tereschenko, G. V.; Goncharov, M. Y.; Soboleva, E. A.; Konopleva, E. I.; Bydanov, O. I.; Plekhov, S. Y.; Grachev, N. S.

2026-06-10 radiology and imaging 10.64898/2026.06.08.26354011 medRxiv
Top 0.1%
28.2%
Show abstract

Background. Lung metastases in osteosarcoma (OS) are the main cause of the death. The accuracy of the diagnosis of nodules by computed tomography (CT) of the lungs is critically important for determining the disseminated stage of the disease and planning surgical treatment. The use of artificial intelligence (AI) in the search for lung nodules increases the accuracy of diagnosis and reduces the chance of missing metastases. Objective: to evaluate the accuracy of lung nodules diagnosis in adolescents with OS using AI. Methods. A retrospective assessment of CT scans of adolescents with OS was performed. A pathological nodule with an average size of [&ge;]4 mm was considered a target finding. The diagnostic accuracy of an AI algorithm previously trained on an adult dataset was evaluated, and the number of false positives (FP) and false negatives (FN) was determined. Sensitivity, specificity, accuracy, area under the ROC curve (AUC), positive predictive value, negative predictive value, and F1-measure were calculated. Based on the obtained results, the effectiveness of the algorithm was assessed. Results. 248 CT scans of adolescents with OS were evaluated. The following results were obtained: in 5 cases, the AI algorithm showed a FP result (2.02%), in 34 cases, it showed a FN result (13.71%), and in 209 cases, a correct result (both true positive and true negative) (84.27%). The diagnostic accuracy of the algorithm was 0.843 (95% CI 0.794-0.887). The application of the AI algorithm in the practice of an X-ray doctor in a specific clinical task would allow to increase the sensitivity from 0.805 to 0.891, while ensuring an absolute decrease in the number of FN results by 8.59% and a relative decrease by 44%. Conclusion. The obtained results confirm the practical value of the application of the AI algorithm and justify the implementation of AI-assisted systems in the diagnostic protocols for lung metastases in adolescents with OS.

6
Automated assessment of chest CT severity scores in patients suspected of COVID-19 infection

Ajmera, P.; Rathi, S.; Dosi, U.; Kalli, S. L.; Luthra, A.; Khaladkar, S.; Pant, R.; Seth, J.; Mishra, P.; Gawali, M.; Pargaonkar, Y.; Kulkarni, V.; Kharat, A.

2022-12-30 radiology and imaging 10.1101/2022.12.28.22284027 medRxiv
Top 0.1%
28.1%
Show abstract

BackgroundThe COVID-19 pandemic has claimed numerous lives in the last three years. With new variants emerging every now and then, the world is still battling with the management of COVID-19. PurposeTo utilize a deep learning model for the automatic detection of severity scores from chest CT scans of COVID-19 patients and compare its diagnostic performance with experienced human readers. MethodsA deep learning model capable of identifying consolidations and ground-glass opacities from the chest CT images of COVID-19 patients was used to provide CT severity scores on a 25-point scale for definitive pathogen diagnosis. The model was tested on a dataset of 469 confirmed COVID-19 cases from a tertiary care hospital. The quantitative diagnostic performance of the model was compared with three experienced human readers. ResultsThe test dataset consisted of 469 CT scans from 292 male (average age: 52.30 {+/-} 15.90 years) and 177 female (average age: 53.47 {+/-} 15.24) patients. The standalone model had an MAE of 3.192, which was lower than the average radiologists MAE of 3.471. The model achieved a precision of 0.69 [0.65, 0.74] and an F1 score of 0.67 [0.62, 0.71], which was significantly superior to the average reader precision of 0.68 [0.65, 0.71] and F1 score of 0.65 [0.63, 0.67]. The model demonstrated a sensitivity of 0.69 [95% CI: 0.65, 0.73] and specificity of 0.83 [95% CI: 0.81, 0.85], which was comparable to the performance of the three human readers, who had an average sensitivity of 0.71 [95% CI: 0.69, 0.73] and specificity of 0.84 [95% CI: 0.83, 0.85]. ConclusionThe AI model provided explainable results and performed at par with human readers in calculating CT severity scores from the chest CT scans of patients affected with COVID-19. The model had a lower MAE than that of the radiologists, indicating that the CTSS calculated by the AI was very close in absolute value to the CTSS determined by the reference standard.

7
Comparison of the diagnostic accuracy among GPT-4 based ChatGPT, GPT-4V based ChatGPT, and radiologists in musculoskeletal radiology

Horiuchi, D.; Tatekawa, H.; Oura, T.; Shimono, T.; Walston, S. L.; Takita, H.; Matsushita, S.; Mitsuyama, Y.; Miki, Y.; Ueda, D.

2023-12-09 radiology and imaging 10.1101/2023.12.07.23299707 medRxiv
Top 0.1%
27.4%
Show abstract

ObjectiveTo compare the diagnostic accuracy of Generative Pre-trained Transformer (GPT)-4 based ChatGPT, GPT-4 with vision (GPT-4V) based ChatGPT, and radiologists in musculoskeletal radiology. Materials and MethodsWe included 106 "Test Yourself" cases from Skeletal Radiology between January 2014 and September 2023. We input the medical history and imaging findings into GPT-4 based ChatGPT and the medical history and images into GPT-4V based ChatGPT, then both generated a diagnosis for each case. Two radiologists (a radiology resident and a board-certified radiologist) independently provided diagnoses for all cases. The diagnostic accuracy rates were determined based on the published ground truth. Chi-square tests were performed to compare the diagnostic accuracy of GPT-4 based ChatGPT, GPT-4V based ChatGPT, and radiologists. ResultsGPT-4 based ChatGPT significantly outperformed GPT-4V based ChatGPT (p < 0.001) with accuracy rates of 43% (46/106) and 8% (9/106), respectively. The radiology resident and the board-certified radiologist achieved accuracy rates of 41% (43/106) and 53% (56/106). The diagnostic accuracy of GPT-4 based ChatGPT was comparable to that of the radiology resident but was lower than that of the board-certified radiologist, although the differences were not significant (p = 0.78 and 0.22, respectively). The diagnostic accuracy of GPT-4V based ChatGPT was significantly lower than those of both radiologists (p < 0.001 and < 0.001, respectively). ConclusionGPT-4 based ChatGPT demonstrated significantly higher diagnostic accuracy than GPT-4V based ChatGPT. While GPT-4 based ChatGPTs diagnostic performance was comparable to radiology residents, it did not reach the performance level of board-certified radiologists in musculoskeletal radiology.

8
Evaluation of an artificial intelligence model for opportunistic calculation of Agatston score on non-gated computed tomography of the chest

McKinney, S. E.; Mercaldo, S. F.; Chin, J. K.; Ghatak, A.; Halle, M. A.; Hedgire, S. S.; Meyersohn, N. M.; Ghoshhajra, B.; Dreyer, K.; Kalra, M. K.; Bizzo, B.; Hillis, J. M.

2024-11-21 radiology and imaging 10.1101/2024.11.20.24317666 medRxiv
Top 0.1%
27.4%
Show abstract

ImportanceThe Agatston score is a measure of cardiovascular disease traditionally calculated on cardiac gated computed tomography (CT) of the chest. Cardiac gated CT is resource-intensive, can be hard to access, and involves extra radiation exposure. Artificial intelligence (AI) can be used to opportunistically calculate Agatston score on non-gated CTs performed for other indications. ObjectiveThis study compared the accuracy of an AI model (Riverain Technologies ClearRead CT CAC) at calculating Agatston scores on non-gated CTs to both consensus radiologist interpretations on the same CTs and Agatston scores from paired cardiac gated CTs. DesignA retrospective standalone performance assessment was conducted on a dataset of non-contrast CT chest cases acquired between January 2022 and December 2023. SettingThe study was conducted at five hospitals in the United States. ParticipantsThe cohort included non-gated CTs from 491 patients. It was enriched to ensure a representation of disease severity by selecting approximately two-thirds of patients using the originally reported Agatston score on a paired cardiac gated CT within the study timeframe. Main Outcome(s) and Measure(s)The study compared the agreement of Agatston categories (0, 1-99, 100-399 and [&ge;]400) between the AI model and ground truth radiologists or original radiology reports using the quadratic weighted Kappa coefficient. Exposure(s)Each non-gated CT case was interpreted independently by three radiologists to establish consensus interpretations. Each CT was then interpreted by the AI model. The Agatston scores for paired cardiac gated CTs were obtained from original radiology reports. ResultsThe agreement between the AI model and ground truth radiologists was 0.959 (95% CI: 0.943-0.975). This result was broadly consistent across sex, age group, race, ethnicity and CT scanner manufacturer subgroups. The agreement between the AI model and paired cardiac gated CT was 0.906 (95% CI: 0.882-0.927). Conclusions and RelevanceThe assessed AI model accurately calculated Agatston scores on non-gated CTs and produced similar scores to paired cardiac gated CTs. Its use could broaden screening for atherosclerotic cardiovascular disease, enabling opportunistic screening on CTs captured for other indications.

9
The Expertise Paradox: Who Benefits from LLM-Assisted Brain MRI Differential Diagnosis?

Schramm, S.; Le Guellec, B.; Topka, M.; Svec, M.; Backhaus, P.; Eisenkolb, V. M.; Riedel, E. O.; Beyrle, M.; Platzek, P.-S.; Ramschütz, C.; Paprottka, K. J.; Renz, M.; Bodden, J.; Kirschke, J. S.; Ziegelmeyer, S.; Busch, F.; Makowski, M. R.; Adams, L. C.; Bressem, K. K.; Hedderich, D. M.; Wiestler, B.; Kim, S. H.

2025-10-28 radiology and imaging 10.1101/2025.10.28.25338816 medRxiv
Top 0.1%
27.3%
Show abstract

PurposeTo evaluate how reader experience influences the diagnostic benefit from LLM assistance in brain MRI differential diagnosis. Materials and MethodsNeuroradiologists (n = 4), radiology residents (n = 4), and neurology/neurosurgery residents (n = 4) were recruited. A dataset of complex brain MRI cases was curated from the local imaging database (n = 40). For each case, readers provided a textual description of the main imaging finding and their top three differential diagnoses ("Unassisted"). Three state-of-the-art large language models (GPT-4.1, Gemini 2.5 Pro, DeepSeek-R1) were prompted to generate top-three differentials based on the clinical case description and reader-specific findings. Readers then revised their differential diagnoses after reviewing GPT-4.1 suggestions ("Assisted"). To evaluate the association between reader experience and diagnostic benefit, a cumulative link mixed model (CLMM) was fitted, with change in diagnostic result as ordinal outcome, reader experience as predictor, and random intercepts for rater and case. ResultsLLM-generated differential diagnoses achieved the highest top-3 accuracy when provided with image descriptions from neuroradiologists (top-3: 78.8-83.8%), followed by radiology residents (top-3: 71.8-77.6%), and neurology/neurosurgery residents (top-3: 62.6-64.5%). In contrast, mean relative gains in top-3 accuracy through LLM assistance diminished with increasing experience, with +19.2% for neurology/neurosurgery residents (from 43.2% to 62.6%), +14.7% for radiology residents (from 59.6% to 74.4%), and +4.4% for neuroradiologists (from 83.1% to 87.5%). The CLMM demonstrated a significant negative association between reader experience and diagnostic benefit from LLM assistance ({beta} = -0.10, p = 0.005). ConclusionWith increasing reader experience, absolute diagnostic LLM performance with reader-generated input improved, while relative diagnostic gains through LLM assistance paradoxically diminished. Our findings call attention to the divergence between standalone LLM performance and clinically relevant reader benefit, and emphasize the need to account for human-AI interaction in this context.

10
Artificial Intelligence-assisted reader evaluation in acute CT head interpretation (AI-REACT): a multireader multicase study.

Novak, A.; Shah, R.; Espinosa Morgado, A. T.; Robert, D.; Kumar, S.; Oke, J.; Bhatia, K.; Romsauerova, A.; Das, T.; Narbone, M.; Dharmadhikari, R.; Harrison, M.; Vimalesvaran, K.; Gooch, J.; Woznitza, N.; Lowe, D. J.; Shuaib, H.; Ather, S.; AI-REACT Reader Study Group,

2025-10-17 radiology and imaging 10.1101/2025.10.17.25337471 medRxiv
Top 0.1%
26.8%
Show abstract

BackgroundNon-contrast CT head scans (NCCTH) are the most frequently requested cross-sectional imaging in the Emergency Department. While AI tools have been developed to detect NCCTH abnormalities, most validation studies compare AI to radiologists, with limited evidence on the impact of AI assistance for other healthcare professionals. ObjectiveTo evaluate whether an AI-powered tool improves the accuracy, speed, and confidence of general radiologists, emergency clinicians, and radiographers in detecting critical abnormalities on NCCTH, and to assess the tools stand-alone performance and factors influencing diagnostic accuracy and efficiency. MethodsA retrospective dataset of 150 NCCTH (52 normal, 98 with critical abnormalities: intracranial haemorrhage, hypodensity, midline shift, mass effect, or skull fracture) was reviewed by 30 readers (10 radiologists, 15 emergency clinicians, 5 radiographers) from four NHS trusts. Each reader interpreted scans first unaided, then with the qER EU 2.0 AI tool, separated by a 2-week washout. Ground truth was established by consensus of two neuroradiologists. We assessed the stand-alone performance of qER and its effect on reader diagnostic accuracy, confidence, and interpretation speed. ResultsThe qER algorithm demonstrated strong diagnostic performance across most pathology subgroups (AUC 0.821-0.976). With AI assistance, pooled reader sensitivity for critically abnormal scans increased from 82.8% to 89.7% (+6.9%, 95% CI +1.4% to +10.6%, p<0.001), and for intracranial haemorrhage from 84.6% to 91.6% (+7.0%, 95% CI +3.2% to +10.8%, p<0.001), but specificity decreased from 84.5% to 78.9% (-5.5%, 95% CI -11.0% to -0.09%, p=0.046). Reader confidence AUC did not change significantly. ED clinicians with AI achieved sensitivity comparable to unaided radiologists, with no significant change in specificity. ConclusionAI-assisted interpretation increased reader sensitivity for critical abnormalities but reduced specificity. Notably, AI assistance enabled ED clinicians to reach diagnostic sensitivity similar to unaided radiologists, supporting the potential for AI to extend the diagnostic capabilities of non-radiologists. Further prospective studies are warranted to confirm these findings in real-world settings. FundingThis study was funded by Qure.ai via an NHSX Award EthicsThe study has been approved by the UK Healthcare Research Authority (IRAS 310995, approved 13/12/2022). The use of anonymised retrospective NCCTH has been authorised by Oxford University Hospitals. Trial registration numberNCT06018545. Research in contextO_ST_ABSWhat is already known on this topicC_ST_ABSAI-derived algorithms for the detection of pathological findings on non-contrast CT head (NCCTH) images have previously demonstrated strong diagnostic performance when used on retrospective datasets. AI-assisted image interpretation using these algorithms has been shown to enhance the diagnostic performance of general and neuro-radiologists in silico. The potential for AI to enhance the performance of less skilled readers who may encounter and be required to act on these images in clinical practice (e.g. non-specialist radiologists, emergency medicine clinicians and radiographers) is as yet untested, however. What this study addsThis large multicase multireader study demonstrates that AI-assisted image interpretation may be used to enhance the in silico diagnostic performance of Emergency Department physicians to a level comparable to that of general radiologists. How this study might affect research, practice or policyThis study raises the possibility that AI-assisted image interpretation could be used to assist non-radiologist clinicians in the safe interpretation of NCCTH scans. Further prospective research is required to test this hypothesis in clinical practice and explore the potential for AI-assisted interpretation to support safe discharge of patients with normal or low-risk scans.

11
Predicting pneumothorax after lung biopsy from pre-operative imaging using a deep convolutional neural network

Jang, S. J.; Pua, B. B.; Shih, G.

2022-03-10 radiology and imaging 10.1101/2022.03.07.22271554 medRxiv
Top 0.1%
24.5%
Show abstract

BackgroundPneumothorax remains one of the most common complications after computed tomography (CT)-guided lung biopsies. Radiographic features including bullae and nodule size are possible markers for post-biopsy pneumothorax. We determine whether a convolutional neural network (CNN) can accurately predict a pneumothorax after lung biopsy based on pre-operative imaging alone. MethodsWith institutional review board approval, we retrospectively evaluated 3,822 patients who underwent a CT-guided lung biopsy between 2011 to 2019. Two image sets were created with CT scout images (1300 patients, 650 pneumothoraces) and chest x-rays (CXR) taken within three months pre-procedure (884 patients, 140 pneumothoraces). Using pre-operative images, CNNs of varying layer depths were trained using transfer learning to predict the development of a pneumothorax post-biopsy. Performance against models were compared using sensitivity analysis and the McNemars test. ResultsThe CNN models trained with CT scout images performed near chance. However, the models performed better with CXR radiographs taken within three months pre-biopsy. For the anterior-posterior view, sensitivity was 0.40, specificity was 0.89, PPV was 0.43, and NPV was 0.87 (AUC = 0.67). For the lateral view, sensitivity was 0.40, specificity was 0.80, PPV was 0.32, and NPV was 0.86 (AUC = 0.65). Increasing CNN layers did not affect performance (p > 0.05). ConclusionChest radiographs taken within three months of lung biopsy may provide important radiographic information for CNNs to assess pneumothorax risk in patients prior to CT-guided lung biopsies. However, more baseline and standardized CXRs before biopsies are necessary to create a robust model for clinical application.

12
Distinguishing GPT-4-generated Radiology Abstracts from Original Abstracts: Performance of Blinded Human Observers and AI Content Detector

Ufuk, F.; Peker, H.; Sagtas, E.; Yagci, A. B.

2023-05-03 radiology and imaging 10.1101/2023.04.28.23289283 medRxiv
Top 0.1%
23.5%
Show abstract

ObjectiveTo determine GPT-4s effectiveness in writing scientific radiology article abstracts and investigate human reviewers and AI Content detectors success in distinguishing these abstracts. Additionally, to determine the similarity scores of abstracts generated by GPT-4 to better understand its ability to create unique text. MethodsThe study collected 250 original articles published between 2021 and 2023 in five radiology journals. The articles were randomly selected, and their abstracts were generated by GPT-4 using a specific prompt. Three experienced academic radiologists independently evaluated the GPT-4 generated and original abstracts to distinguish them as original or generated by GPT-4. All abstracts were also uploaded to an AI Content Detector and plagiarism detector to calculate similarity scores. Statistical analysis was performed to determine discrimination performance and similarity scores. ResultsOut of 134 GPT-4 generated abstracts, average of 75 (56%) were detected by reviewers, and average of 50 (43%) original abstracts were falsely categorized as GPT-4 generated abstracts by reviewers. The sensitivity, specificity, accuracy, PPV, and NPV of observers in distinguishing GPT-4 written abstracts ranged from 51.5% to 55.6%, 56.1% to 70%, 54.8% to 60.8%, 41.2% to 76.7%, and 47% to 62.7%, respectively. No significant difference was observed between observers in discrimination performance. ConclusionGPT-4 can generate convincing scientific radiology article abstracts. However, human reviewers and AI Content detectors have difficulty in distinguishing GPT-4 generated abstracts from original ones.

13
Comparison of the Diagnostic Performance from Patient's Medical History and Imaging Findings between GPT-4 based ChatGPT and Radiologists in Challenging Neuroradiology Cases

Horiuchi, D.; Tatekawa, H.; Oura, T.; Oue, S.; Walston, S. L.; Takita, H.; Matsushita, S.; Mitsuyama, Y.; Shimono, T.; Miki, Y.; Ueda, D.

2023-08-29 radiology and imaging 10.1101/2023.08.28.23294607 medRxiv
Top 0.1%
23.5%
Show abstract

PurposeTo compare the diagnostic performance between Chat Generative Pre-trained Transformer (ChatGPT), based on the GPT-4 architecture, and radiologists from patients medical history and imaging findings in challenging neuroradiology cases. MethodsWe collected 30 consecutive "Freiburg Neuropathology Case Conference" cases from the journal Clinical Neuroradiology between March 2016 and June 2023. GPT-4 based ChatGPT generated diagnoses from the patients provided medical history and imaging findings for each case, and the diagnostic accuracy rate was determined based on the published ground truth. Three radiologists with different levels of experience (2, 4, and 7 years of experience, respectively) independently reviewed all the cases based on the patients provided medical history and imaging findings, and the diagnostic accuracy rates were evaluated. The Chi-square tests were performed to compare the diagnostic accuracy rates between ChatGPT and each radiologist. ResultsChatGPT achieved an accuracy rate of 23% (7/30 cases). Radiologists achieved the following accuracy rates: a junior radiology resident had 27% (8/30) accuracy, a senior radiology resident had 30% (9/30) accuracy, and a board-certified radiologist had 47% (14/30) accuracy. ChatGPTs diagnostic accuracy rate was lower than that of each radiologist, although the difference was not significant (p = 0.99, 0.77, and 0.10, respectively). ConclusionThe diagnostic performance of GPT-4 based ChatGPT did not reach the performance level of either junior/senior radiology residents or board-certified radiologists in challenging neuroradiology cases. While ChatGPT holds great promise in the field of neuroradiology, radiologists should be aware of its current performance and limitations for optimal utilization.

14
A Deep Learning Based Automated Detection of Mucus Plugs in Chest CT

Sonoda, Y.; Fukuda, K.; Matsuzaki, H.; Yamagishi, Y.; Miki, S.; Nomura, Y.; Mikami, Y.; Yoshikawa, T.; Hanaoka, S.; Kage, H.; Abe, O.

2025-03-11 radiology and imaging 10.1101/2025.03.09.25323501 medRxiv
Top 0.1%
23.2%
Show abstract

This study presents a novel two stage deep learning algorithm for automated detection of mucus plugs in CT scans of patients with respiratory diseases. Despite the clinical significance of mucus plugs in COPD and asthma where they indicate hypoxemia, reduced exercise tolerance, and poorer outcomes, they remain under evaluated in clinical practice due to labor intensive manual annotation. The developed algorithm first segments both patent and obstructed airways using a VNet-based model pre-trained on normal airway structures and fine-tuned on mucus containing scans. Subsequently, a rule-based post processing method identifies mucus plugs by evaluating cross sectional areas along airway centerlines. Validation on an in-house dataset of 33 CT scans from patients with asthma/COPD demonstrated high sensitivity (93.8%) though modest positive predictive value (18.8%). Performance on an external dataset (LIDC-IDRI) achieved 82.8% sensitivity with 23.5% PPV. While challenges remain in reducing false positives, this automated detection tool shows promise for screening applications in both clinical and research settings, potentially addressing the current gap in mucus plug evaluation within standard practice.

15
Automated AI-Based Ventricular Subcompartment Segmentation and Volumetry in Idiopathic Normal Pressure Hydrocephalus

Mutke, M. A.; Griot, S. A.; Wasserthal, J.; Indrakanti, A. K.; Vishwanathan, N.; Mahmutoglu, M. A.; D'Antonoli, T. A.; Bach, M.; Psychogios, M. N.; Lieb, J. M.

2026-06-15 radiology and imaging 10.64898/2026.06.14.26355627 medRxiv
Top 0.1%
23.2%
Show abstract

Purpose In idiopathic normal pressure hydrocephalus (iNPH), longitudinal monitoring of ventricular size is important for diagnosis and treatment follow-up. This study aimed to validate a fully automated AI model for CT ventricular volumetry with subcompartments and to compare AI-derived volume changes with routine radiology assessments. Methods This retrospective, single-center study included 88 patients with iNPH and 456 non-contrast-enhanced head CT examinations. The model was trained on 38 manually labeled CT scans with 12 ventricular subcompartments. Outcomes included segmentation accuracy, correspondence between AI-derived longitudinal ventricular volume changes and radiology report categories (decreased, unchanged, increased), radiologist detection thresholds for ventricular change, and paired pre- and postoperative volume changes in 22 patients with ventriculoperitoneal shunt. Results Mean segmentation accuracy was high (Dice, 0.83). 91% of 100 segmentations were rated as excellent by an expert neuroradiologist. AI-derived ventricular volume changes corresponded well to radiology report categories (median total ventricular volume changes of -17% in cases reported as decreased, 0% in unchanged cases, and +22% in increased cases; all p < 0.001). Radiologists reported ventricular volume change in 50% of cases at an AI-measured relative volume change of +/-6%, and in 90% of cases at +21% for enlargement and -18% for decrease. After shunt placement, ventricular volume decreased by -8% (median), with the largest relative reductions observed in the right temporal and occipital horns. Conclusions Automated AI-based ventricular segmentation on CT enables accurate and reproducible assessment of ventricular volume changes in iNPH and complements routine radiological evaluation for longitudinal and postoperative monitoring.

16
SISTR: Sinus and Inferior alveolar nerve Segmentation with Targeted Refinement on Cone Beam Computed Tomography images

Misrachi, L.; Covili, E.; Mayard, H.; Alaka, C.; Rousseau, J.; Au, W.

2024-02-18 dentistry and oral medicine 10.1101/2024.02.17.24301683 medRxiv
Top 0.1%
23.2%
Show abstract

BackgroundAccurate delineation of the maxillary sinus and inferior alveolar nerve (IAN) is crucial in dental implantology to prevent surgical complications. Manual segmentation from CBCT scans is labor-intensive and error-prone. MethodsWe introduce SISTR (Sinus and IAN Segmentation with Targeted Refinement), a deep learning framework for automated, high-resolution instance segmentation of oral cavity anatomies. SISTR operates in two stages: first, it predicts coarse segmentation and offset maps to anatomical regions, followed by clustering to identify region centroids. Subvolumes of individual anatomical instances are then extracted and processed by the model for fine structure segmentation. Our model was developed on the most diverse dataset to date for sinus and IAN segmentation, sourced from 11 dental clinics and 10 manufacturers (358 CBCTs for sinus, 499 for IAN). ResultsSISTR shows robust generalizability. It achieves strong segmentation performance on an external test set (98 sinus, 91 IAN CBCTs), reaching average DICE scores of 96.64% (95.38-97.60) for sinus and 83.43% (80.96-85.63) for IAN, representing a significant 10 percentage point improvement in Dice score for IAN compared to single-stage methods. Chamfer distances of 0.38 (0.24-0.60) mm for sinus and 0.88 (0.58-1.27) mm for IAN confirm its accuracy. Its inference time of 4 seconds per scan reduces time required for manual segmentation, which can take up to 28 minutes. ConclusionsSISTR offers a fast, accurate, and efficient solution for the segmentation of critical anatomies in dental implantology, making it a valuable tool in digital dentistry. Plain text summaryAccurately determining the locations of important structures such as the maxillary sinus and inferior alveolar nerve is crucial in dental implant surgery to avoid complications. The conventional method of manually mapping these areas from CBCT scans is time-consuming and prone to errors. To address this issue, we have developed SISTR, an AI-based framework that efficiently and accurately automates this process, trained on extensive datasets, sourced from 11 dental clinics and 10 manufacturers. It surpasses conventional methods by identifying anatomical regions within seconds. SISTR provides a rapid and accurate solution for high-resolution segmentation of critical anatomies in dental implantology, making it a valuable tool in digital dentistry.

17
Total and Stroke Related Imaging Utilization Patterns During the COVID-19 Pandemic

Tu, L. H.; Sharma, R.; Malhotra, A.; Schindler, J. L.; Forman, H. P.

2020-05-26 radiology and imaging 10.1101/2020.05.20.20078915 medRxiv
Top 0.1%
22.9%
Show abstract

During the COVID-19 pandemic, radiology practices are reporting a decrease in imaging volumes. We review total imaging volume, CTA head and neck volume, critical results rate, and stroke intervention rates before and during the COVID-19 pandemic. Total imaging volume as well as CTA head and neck imaging fell approximately 60% since the beginning of the pandemic. Critical results fell 60-70% for total imaging as well as for CTA head and neck. Compared to the same time frame a year prior, the number of stroke codes at the early impact of the pandemic had decreased approximately 50%. Proportional reductions in total imaging volume, stroke-related imaging, and associated critical result reports during the COVID-19 pandemic raise concern for missed stroke diagnoses in our population.

18
Emphysema Quantification and Severity Classification with 3-Dimensional Averaging Kernel and Airways Removal

Zhang, J.; Chaudhari, G. R.; Bondarenko, M.; Sohn, J. H.

2022-11-02 radiology and imaging 10.1101/2022.10.31.22281562 medRxiv
Top 0.1%
22.7%
Show abstract

BackgroundEmphysema is a common pulmonary pathology known to be associated with increased risk of lung cancer and lung biopsy complications. Prevailing quantitation method of calculating voxel-wise percentage of low attenuation area (LAA) of lung tissue from CT scans is prone to noise and error due overcounting of single voxel LAA and incomplete segmentation of airways. PurposeWe aim to develop an accurate algorithm to quantitatively measure emphysema and classify its severity.. Methods and MaterialsTwo chest CT datasets were obtained from two tertiary hospitals as training and external validation datasets. Exclusion criteria included any patients whose emphysema extent was not specified by the accompanying report. The training dataset included 722 patients, and the validation dataset included 1006 patients. Following lung segmentation and airways removal, we applied convolution of the segmented lung with averaging kernels of different sizes in 2D and 3D. Cutoffs between "none," "mild to moderate," and "severe" emphysema were determined via weighted logistic regression on the training dataset, and the categorical emphysema extent was obtained for each patient. The main measure for evaluating model performance was area under the curve (AUC) of the receiver operating characteristic (ROC) on the training dataset and accuracy of classification on both the training and the validation dataset. The 1x1x1 kernel, which is equivalent to the traditional LAA score, was used for comparison to other kernels for performance evaluation. ResultsThe best model used a 3D 3x3x3 kernel for average filtering with airways post processing and achieved a mean AUC of 0.782 and 0.985 for "none"-versus-rest and "severe"-versus-rest classifications respectively. It achieved a 0.676 and 0.757 multiclass classification accuracy on the training and validation dataset respectively. Conclusions and RelevanceWe present an automated pipeline that can achieve accurate emphysema quantification and severity classification. We showed that convolving the segmented lung with a 3D 3x3x3 kernel and post-processing to remove airways can reliably quantify emphysema.

19
Prediction of Subsolid Pulmonary Nodule Evolution from Baseline CT Using Temporal Imaging Models

Bondarenko, M.; Qi, K.; Nowroozi, A.; Kim, J.; Kunzang, B.; Lee, A.; Liu, J.; Tran, N.; Weng, S.; Vella, M.; Chaudhari, G.; Schnizler, T.; Innanje, A.; Chen, T.; Sohn, J. H.

2026-08-13 radiology and imaging 10.64898/2026.08.12.26360292 medRxiv
Top 0.1%
22.7%
Show abstract

Background: Prediction of subsolid pulmonary nodule (SSN) progression from baseline CT may improve risk stratification and surveillance planning, but prior approaches have largely relied on fixed follow-up intervals. Methods: This retrospective single-center study evaluated interval-aware temporal imaging models for predicting future SSN growth and morphology across heterogeneous surveillance durations. A total of 24,946 longitudinal scan pairings derived from 2,543 clinician-reviewed SSNs in 426 patients were analyzed. A discriminative deep learning model predicted interval growth from baseline CT, segmentation masks, and interscan interval information, while a temporally conditioned generative model predicted future lesion morphology. Results: The discriminative model achieved an area under the receiver operating characteristic curve of 0.772 (95% confidence interval: 0.704-0.818), with sensitivity of 80.2% and specificity of 58.7% on the test cohort. The generative model predicted future lesion morphology with a Dice similarity coefficient of 0.706 +/-0.186. Prediction performance decreased with increasing follow-up duration, although both models generalized across intervals ranging from months to years. Conclusion: Interval-aware temporal imaging models enable the prediction of future SSN growth and morphology from baseline CT while accounting for variable surveillance intervals. These findings suggest a framework for time-aware, personalized risk assessment that may support individualized surveillance strategies and future AI-assisted management of pulmonary adenocarcinoma spectrum lesions.

20
Impact of Multimodal Prompt Elements on Diagnostic Performance of GPT-4(V) in Challenging Brain MRI Cases

Schramm, S.; Preis, S.; Metz, M.-C.; Jung, K.; Schmitz-Koep, B.; Zimmer, C.; Wiestler, B.; Hedderich, D. M.; Kim, S. H.

2024-03-06 radiology and imaging 10.1101/2024.03.05.24303767 medRxiv
Top 0.1%
19.8%
Show abstract

BackgroundRecent studies have explored the application of multimodal large language models (LLMs) in radiological differential diagnosis. Yet, how different multimodal input combinations affect diagnostic performance is not well understood. PurposeTo evaluate the impact of varying multimodal input elements on the accuracy of GPT-4(V)-based brain MRI differential diagnosis. MethodsThirty brain MRI cases with a challenging yet verified diagnosis were selected. Seven prompt groups with variations of four input elements (image, image annotation, medical history, image description) were defined. For each MRI case and prompt group, three identical queries were performed using an LLM-based search engine ((C) PerplexityAI, powered by GPT-4(V)). Accuracy of LLM-generated differential diagnoses was rated using a binary and a numeric scoring system and analyzed using a chi-square test and a Kruskal-Wallis test. Results were corrected for false discovery rate employing the Benjamini-Hochberg procedure. Regression analyses were performed to determine the contribution of each individual input element to diagnostic performance. ResultsThe prompt group containing an annotated image, medical history, and image description as input exhibited the highest diagnostic accuracy (67.8% correct responses). Significant differences were observed between prompt groups, especially between groups that contained the image description among their inputs, and those that did not. Regression analyses confirmed a large positive effect of the image description on diagnostic accuracy (p << 0.001), as well as a moderate positive effect of the medical history (p < 0.001). The presence of unannotated or annotated images had only minor or insignificant effects on diagnostic accuracy. ConclusionThe textual description of radiological image findings was identified as the strongest contributor to performance of GPT-4(V) in brain MRI differential diagnosis, followed by the medical history. The unannotated or annotated image alone yielded very low diagnostic performance. These findings offer guidance on the effective utilization of multimodal LLMs in clinical practice.