The Presence and Nature of AI-Use Disclosure Statements in Medical Education Journals: A Bibliometric Study
Ans, M.; Maggio, L.; Algodi, H.; Costello, J.; Driessen, E.; Oswald, K.; Lingard, L.
Show abstract
BackgroundAs AI-use becomes more common in research, disclosure policies have emerged to ensure transparency and appropriateness. However, database research in other fields suggests that disclosure may lag behind AI-use. Medical education journal editors report that submitted manuscripts rarely include AI-use disclosures, and they perceive a lack of clarity regarding when and how AI-use should be disclosed. However, we lack objective evidence regarding the incidence and nature of AI-use disclosure in medical education. MethodsUsing bibliometric methods, we searched a database of 24 leading medical education journals for articles published between January and July 2025 (n=2,762 articles). Screening with Covidence software excluded 716 non-empirical and/or non-English language articles. The remainder (n=2,046) were examined for the presence of AI-use disclosures, which were content-analyzed. Results2.5% of empirical articles (n=51) had an AI disclosure statement. BMC Medical Education contained the most disclosures (24), followed by Medical Teacher (7) and Journal of Surgical Education (4). Forty-two articles were authored in non-native English-speaking countries, and 69.4% of all first authors had begun publishing in the past decade. Disclosures averaged 43 words and described use superficially: most commonly "editing" and "translation". Of 18 named tools, ChatGPT was most common. Most disclosures explicitly attested to author responsibility for AI-produced material. Disclosures usually appeared in acknowledgements; those located in methods lacked responsibility attestation. Negative disclosures attesting that AI was not used were also present. DiscussionAI-use disclosures in medical education journals are rare and appear mostly in work from non-native English-speaking regions of the world. A shared disclosure practice is evident: name the tool and affirm author responsibility, but describe use superficially. This suggests a practice of "safe" disclosure that may be more performative than informative, therefore failing to satisfy the goal of ensuring transparent and ethical AI use in research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Knowledge and motivations of training in peer review: an international cross-sectional survey 97%
- Transparency in peer review: Exploring the content and tone of reviewers' confidential comments to editors 96%
- Science in motion: A qualitative analysis of journalists' use and perception of preprints 95%
Similar papers in this journal
- Publishing at any cost: a cross-sectional study of the amount that medical researchers spend on open-access publishing each year 96%
- Comparison of preprints and final journal publications from COVID-19 Studies: Discrepancies in results reporting and spin in interpretation 94%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 93%
Similar papers in this journal
Similar papers in this journal
- A Systematic Examination of Generative Artificial Intelligence (GAI) Usage Guidelines for Scholarly Publishing in Medical Journals 97%
- Checklists to Detect Potential Predatory Biomedical Journals: A Systematic Review 95%
- Evidence of Unreliable Data and Poor Data Provenance in Clinical Prediction Model Research and Clinical Practice 91%
Similar papers in this journal
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 94%
- Citation tracking for systematic literature searching: a scoping review 94%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.