Comparing AI and Human Coding of NIH Grant Abstracts to Identify Innovations in Opioid Addiction Treatment
Alkhatib, S. A.; Jiwa, N.; Judd, D.; Luningham, J. M.; Sawyer-Morris, G.; Ulukaya, M.; Molfenter, T.; Taxman, F. S.; Walters, S. T.
Show abstract
Large language models (LLMs) are increasingly used for qualitative analysis in substance use research, yet their performance relative to human coders remains underexplored. This study compares ChatGPT-4.0 with human coders in identifying and describing the core innovation of NIH grants focused on reducing opioid overdose. A total of 118 NIH HEAL Initiative grant abstracts were independently coded by ChatGPT and humans to generate innovation descriptions, which were then evaluated by both human raters and ChatGPT for depth/detail and relevance/completeness using 5-point Likert scales. Identical instructions were used across all coding and evaluation stages. ChatGPT-generated descriptions were consistently rated higher than human-generated descriptions on both dimensions. Human evaluators rated ChatGPT outputs at an average of 4.47 for both depth/detail and relevance/completeness, compared to 3.33 and 3.24 for human outputs, respectively (F(1,176)=133.9, p<0.001). These findings suggest that LLMs, when carefully prompted, can enhance the efficiency and quality of qualitative research evaluation.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 95%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
Similar papers in this journal
Similar papers in this journal
- Dissemination and Stakeholder Engagement Practices Among Dissemination & Implementation Scientists: Results from an Online Survey 95%
- Stakeholders’ views on an institutional dashboard with metrics for responsible research 94%
- Clinical code sets and the problem of redundancy in code set repositories 94%
Similar papers in this journal
- Scientific hypothesis generation process in clinical research: a secondary data analytic tool versus experience study protocol 94%
- Strategies and Tools for electronic health records and physician workflow alignment: A scoping review protocol 92%
- The PUPPY Study - Protocol for a Longitudinal Mixed Methods Study Exploring Problems Coordinating and Accessing Primary Care for Attached and Unattached Patients Exacerbated During the COVID-19 Pandemic Year 90%
Similar papers in this journal
- Data-driven hypothesis generation among junior clinical researchers: A comparison of a secondary data analysis with visualization (VIADS) and other tools 94%
- Implementation and Impact of a Diversity Supplement Repository 93%
- Emotional Distress, Stress, Anxiety and the Impact of the COVID-19 Pandemic on Early Career Women in Healthcare Sciences Research 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.