Distinguishing GPT-4-generated Radiology Abstracts from Original Abstracts: Performance of Blinded Human Observers and AI Content Detector
Ufuk, F.; Peker, H.; Sagtas, E.; Yagci, A. B.
Show abstract
ObjectiveTo determine GPT-4s effectiveness in writing scientific radiology article abstracts and investigate human reviewers and AI Content detectors success in distinguishing these abstracts. Additionally, to determine the similarity scores of abstracts generated by GPT-4 to better understand its ability to create unique text. MethodsThe study collected 250 original articles published between 2021 and 2023 in five radiology journals. The articles were randomly selected, and their abstracts were generated by GPT-4 using a specific prompt. Three experienced academic radiologists independently evaluated the GPT-4 generated and original abstracts to distinguish them as original or generated by GPT-4. All abstracts were also uploaded to an AI Content Detector and plagiarism detector to calculate similarity scores. Statistical analysis was performed to determine discrimination performance and similarity scores. ResultsOut of 134 GPT-4 generated abstracts, average of 75 (56%) were detected by reviewers, and average of 50 (43%) original abstracts were falsely categorized as GPT-4 generated abstracts by reviewers. The sensitivity, specificity, accuracy, PPV, and NPV of observers in distinguishing GPT-4 written abstracts ranged from 51.5% to 55.6%, 56.1% to 70%, 54.8% to 60.8%, 41.2% to 76.7%, and 47% to 62.7%, respectively. No significant difference was observed between observers in discrimination performance. ConclusionGPT-4 can generate convincing scientific radiology article abstracts. However, human reviewers and AI Content detectors have difficulty in distinguishing GPT-4 generated abstracts from original ones.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 96%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
- Predicting EGFR mutation status in lung adenocarcinoma presenting as ground-glass opacity: utilizing radiomics model in clinical translation 92%
Similar papers in this journal
- Investigating the relationship between internal spinal alignment and back shape in patients with scoliosis using PCdare: a comparative, reliability and validation study 93%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 93%
- Accuracy of deep learning based computed tomography diagnostic system of COVID-19: a consecutive sampling external validation cohort study 92%
Similar papers in this journal
- “This is a quiz” Premise Input: A Key to Unlocking Higher Diagnostic Accuracy in Large Language Models 94%
- Effects of contrast-medium and vertebral measurement level on computed tomography-based body composition parameters of skeletal muscle and adipose tissue 93%
- Benchmarking Deep Learning-based Image Retrieval of Oral Tumor Histology 91%
Similar papers in this journal
- Computed Tomography Features of COVID-19 in Children: A Systematic Review and Meta-analysis 89%
- The PBL teaching method in Neurology Education in the Traditional Chinese Medicine undergraduate students: An Observational Study 89%
- Open Science Practices Among Authors Published in Complementary, Alternative, and Integrative Medicine Journals: An International, Cross-Sectional Survey 89%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 96%
- Inconsistency of AI in Intracranial Aneurysm Detection with Varying Dose and Image Reconstruction 92%
- Evaluation of the second-generation whole-heart motion correction algorithm (SSF2) used to demonstrate the aortic annulus on cardiac CT 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.