Evaluative Stance Toward Artificial Intelligence in High-Quartile Medical Journals (2021-2026): Large-Scale LLM-Assisted Computational Content Analysis
Wang, L.; Poenaru, D.
Show abstract
Background: Medical-AI publications do more than report technical performance; they also frame AI as beneficial, uncertain, or risky. How this evaluative stance has changed across the medical literature is not well characterized. Objective: To characterize evaluative stance in published medical-AI discourse abstracts from January 2021 through April 2026 and examine variation over time, concern themes, failure mechanisms, specialties, first-author geography, and publication format. Methods: We conducted an LLM-assisted computational content analysis of medical-AI abstracts from first- and second-quartile medical journals. Of 97,492 post-cutoff Q1/Q2 records entering the prefilter, 16,759 were retained as discourse or evaluative. Claude Sonnet 4.6 assigned 16,749 valid stance classifications using Alarm, Caution, Neutral, Cautious Optimism, and Advocacy. Annual analyses used 16,747 records dated 2021-2026. Critical stance was defined as Alarm plus Caution and indexed evaluative scrutiny rather than opposition or author psychology. Each LLM step was validated against blinded human coding by one author: prefilter Cohen kappa = 0.51, stance quadratic-weighted kappa = 0.79 (95% CI 0.72-0.84) for codable, in-scope records, specialty kappa = 0.75, and mechanism-axis kappa = 0.84 for model type and 0.57 for failure mode. Results: Advocacy declined from 2.9% in 2021 to 0.6% in partial 2026, while Cautious Optimism remained the majority stance. Among 16,749 valid classifications, 30.8% were critical. Critical share increased from 25.4% to 32.6%, a 7.25-percentage-point increase based on unrounded estimates. Among critical records, patient safety remained the most prevalent concern. Hallucination/errors increased by 30.9 percentage points. Regulation declined by 22.0 percentage points and ethics/bias by 8.1 percentage points in prevalence share; these declines do not necessarily indicate lower publication counts. Within the hallucination/error theme, factual error was more common than fabrication. Fabrication estimates should be treated as an upper bound because failure-mode agreement was moderate. Specialty patterns were heterogeneous. Critical rate was inversely associated with FDA-cleared device availability (Spearman rho = -0.65, two-sided p = 0.004), which does not measure adoption, deployment, maturity, or clinical use. First-author geography described publication metadata and discourse, not national attitudes or research quality. Reviews were the least critical and most favourable format. In exploratory forward validation, 2 of 78 early Advocacy predictions were fully borne out, although the analysis was single-rater and retrieval-dependent. Conclusions: Published medical-AI abstracts became modestly less promotional and more focused on specific errors and safety concerns. Unqualified promotion declined, but qualified favourable framing remained dominant, and the rise in critical stance was modest. Concern moved toward errors and patient safety, with factual error discussed more often than fabrication. These findings describe published discourse, not AI capability or whether the evaluations were correct.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 94%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 92%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 91%
Similar papers in this journal
Similar papers in this journal
- Quantifying new threats to health and biomedical literature integrity from rapidly scaled publications and problematic research 91%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 91%
- High-cited favorable studies for COVID-19 treatments ineffective in large trials 90%
Similar papers in this journal
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 90%
- A Crowdsourcing Approach to Develop Machine Learning Models to Quantify Radiographic Joint Damage in Rheumatoid Arthritis 89%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.