Back

Evaluative Stance Toward Artificial Intelligence in High-Quartile Medical Journals (2021-2026): Large-Scale LLM-Assisted Computational Content Analysis

Wang, L.; Poenaru, D.

2026-07-27 health informatics
10.64898/2026.07.23.26358815 medRxiv
Show abstract

Background: Medical-AI publications do more than report technical performance; they also frame AI as beneficial, uncertain, or risky. How this evaluative stance has changed across the medical literature is not well characterized. Objective: To characterize evaluative stance in published medical-AI discourse abstracts from January 2021 through April 2026 and examine variation over time, concern themes, failure mechanisms, specialties, first-author geography, and publication format. Methods: We conducted an LLM-assisted computational content analysis of medical-AI abstracts from first- and second-quartile medical journals. Of 97,492 post-cutoff Q1/Q2 records entering the prefilter, 16,759 were retained as discourse or evaluative. Claude Sonnet 4.6 assigned 16,749 valid stance classifications using Alarm, Caution, Neutral, Cautious Optimism, and Advocacy. Annual analyses used 16,747 records dated 2021-2026. Critical stance was defined as Alarm plus Caution and indexed evaluative scrutiny rather than opposition or author psychology. Each LLM step was validated against blinded human coding by one author: prefilter Cohen kappa = 0.51, stance quadratic-weighted kappa = 0.79 (95% CI 0.72-0.84) for codable, in-scope records, specialty kappa = 0.75, and mechanism-axis kappa = 0.84 for model type and 0.57 for failure mode. Results: Advocacy declined from 2.9% in 2021 to 0.6% in partial 2026, while Cautious Optimism remained the majority stance. Among 16,749 valid classifications, 30.8% were critical. Critical share increased from 25.4% to 32.6%, a 7.25-percentage-point increase based on unrounded estimates. Among critical records, patient safety remained the most prevalent concern. Hallucination/errors increased by 30.9 percentage points. Regulation declined by 22.0 percentage points and ethics/bias by 8.1 percentage points in prevalence share; these declines do not necessarily indicate lower publication counts. Within the hallucination/error theme, factual error was more common than fabrication. Fabrication estimates should be treated as an upper bound because failure-mode agreement was moderate. Specialty patterns were heterogeneous. Critical rate was inversely associated with FDA-cleared device availability (Spearman rho = -0.65, two-sided p = 0.004), which does not measure adoption, deployment, maturity, or clinical use. First-author geography described publication metadata and discourse, not national attitudes or research quality. Reviews were the least critical and most favourable format. In exploratory forward validation, 2 of 78 early Advocacy predictions were fully borne out, although the analysis was single-rater and retrieval-dependent. Conclusions: Published medical-AI abstracts became modestly less promotional and more focused on specific errors and safety concerns. Unqualified promotion declined, but qualified favourable framing remained dominant, and the rise in critical stance was modest. Concern moved toward errors and patient safety, with factual error discussed more often than fabrication. These findings describe published discourse, not AI capability or whether the evaluations were correct.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.