Citation reliability of frontier large language models in medical writing and its automated verification
Shin, R.; Lee, J.-M.; Park, J.; Kwun, J.-S.; Cho, H.-W.; Kang, S.-H.; Jeon, K.-H.
Show abstract
Large language models (LLMs) are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them. We evaluated three frontier LLMs, Claude Opus 4.8, GPT-5.5, and Gemini 3.5 Flash, generating 270 cardiology narrative reviews with web search enabled, and verified all 8,050 references against PubMed. Problematic references accounted for 11.5% of GPT-5.5 output, 29.2% of Claude output, and 29.6% of Gemini output (P < 0.001), with no significant gradient across topics of differing publication volume (P = 0.052). Misattribution, a valid PubMed identifier that resolves to a different article, made up 77% of errors, whereas fabrication was rare (0.6%). Against an expert-adjudicated set of 270 references, an LLM-based Chain-of-Verification (CoVe) detected 60 of 62 problematic references (sensitivity 96.8%, specificity 98.6%), including every misattribution and fabrication. LLM-generated citations require identifier-level verification, and CoVe provides it at expert-level accuracy.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 95%
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 91%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 91%
Similar papers in this journal
- Large-scale validation of the Prediction model Risk Of Bias ASsessment Tool (PROBAST) using a short form: high risk of bias models show poorer discrimination 92%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 92%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 91%
Similar papers in this journal
- Toward Assessing Clinical Trial Publications for Reporting Transparency 93%
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 92%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 91%
Similar papers in this journal
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 92%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 92%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.