A Large Language Model-based Approach for Analyzing Covariates of Health Equity in Registered Research Projects
Nananukul, N.; Kejriwal, M.
Show abstract
Large language models (LLMs) have made significant advancements in natural language processing, offering broad applications in multiple domains. This study explores the use of the GPT-3.5 LLM to conduct efficient and robust computational analysis of registered research projects on the All of Us platform. Specifically, we explore the association between projects pursuing health equity research and: the projects use of demographic categories (which All of Us enables), the multi-institutional composition of the team leading the project, and the involvement of R2 institutions (compared to only R1 institutions). We demonstrate the utility of GPT-3.5 in automating tasks ranging from generating Python scripts for extracting attributes from free text (such as project description and goals) to identifying and classifying institutions as R1 and R2, and summarizing project details into Unified Medical Language System (UMLS)-coded medical keywords. These contributions significantly reduced manual workload, allowing researchers to focus on more in-depth analysis. Our results reveal health equity insights not readily available in the original All of Us research hub. Specifically, we find a strong positive association between the use of demographic data and projects focused on health equity, while other associations such as health equity projects conducted by institutions were positive but weaker and more dependent on specific project topics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Diversity and inclusion: A hidden additional benefit of Open Data 96%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
Similar papers in this journal
- The Pandemic Journaling Project: A new dataset of first-person accounts of the COVID-19 pandemic 95%
- Analysis of the Health Economics Portfolio Funded by the National Institutes of Health in Response to Published Guidance 93%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 93%
Similar papers in this journal
- Identifying and preventing fraudulent responses in online public health surveys: Lessons learned during the COVID-19 pandemic 92%
- Architecture of systems affecting disease trajectories in a conflict zone: A community-centered systems inquiry in North Gaza 91%
- An emerging TIMER-2C framework for addressing barriers to research culture and productivity among local healthcare providers in the Middle East and sub-Saharan Africa: a qualitative study and modified Delphi approach 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.