Back

Comparing human and artificial intelligence in writing for health journals: an exploratory study

Haq, Z.; Naeem, H.; Naeem, A.; Iqbal, F.; Zaeem, D.

2023-02-26 public and global health
10.1101/2023.02.22.23286322 medRxiv
Show abstract

Aim and objectivesThe aim was to contribute to the editorial principles on the possible use of Artificial Intelligence (AI)- based tools for scientific writing. The objectives included O_LIEnlist the inclusion and exclusion criteria to test ChatGPT use in scientific writing C_LIO_LIDevelop evaluation criteria to assess the quality of articles written by human authors and ChatGPT C_LIO_LICompare prospectively written manuscripts by human authors and ChatGPT C_LI DesignProspective exploratory study InterventionHuman authors and ChatGPT were asked to write short journal articles on three topics: 1) Promotion of early childhood development in Pakistan 2) Interventions to improve gender-responsive health services in low-and-middle-income countries, and 3) The pitfalls in risk communication for COVID-19. We content analyzed the articles using an evaluation matrix. Outcome measuresThe completeness, credibility, and scientific content of an article. Completeness meant that structure (IMRaD) and organization was maintained. Credibility required that others work is duly cited, with an accurate bibliography. Scientific content required specificity, data accuracy, cohesion, inclusivity, confidentiality, limitations, readability, and time efficiency. ResultsThe articles by human authors scored better than ChatGPT in completeness and credibility. Similarly, human-written articles scored better for most of the items in scientific content except for time efficiency where ChatGPT scored better. The methods section was absent in ChatGPT articles, and a majority of references in its bibliography were unverifiable. ConclusionsChatGPT generates content that is believable but may not be true. The creators of this powerful model must step up and provide solutions to manage its glitches and potential misuse. In parallel, the academic departments, editors, and publishers must expect a growing utilization of ChatGPT and similar tools. Disallowing ChatGPT as a co-author may not be enough on their part. They must adapt the editorial policies, use measures to detect AI-based writing, and stop its likely implications for human health and life. STRENGTHS AND LIMITATIONSO_LIFirst study that examines the scientific writing of ChatGPT by comparing it with human-written articles. C_LIO_LIExplains how ChatGPT generates believable content that may not be true. C_LIO_LIIndicates that the creators of ChatGPT must step up to address its misuse and potential hazards. C_LIO_LIAn initial exploration, based on limited data--larger studies are needed for generalizable conclusions. C_LI

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.