Back

A Large Language Model-based Approach for Analyzing Covariates of Health Equity in Registered Research Projects

Nananukul, N.; Kejriwal, M.

2024-09-26 public and global health
10.1101/2024.09.24.24314327 medRxiv
Show abstract

Large language models (LLMs) have made significant advancements in natural language processing, offering broad applications in multiple domains. This study explores the use of the GPT-3.5 LLM to conduct efficient and robust computational analysis of registered research projects on the All of Us platform. Specifically, we explore the association between projects pursuing health equity research and: the projects use of demographic categories (which All of Us enables), the multi-institutional composition of the team leading the project, and the involvement of R2 institutions (compared to only R1 institutions). We demonstrate the utility of GPT-3.5 in automating tasks ranging from generating Python scripts for extracting attributes from free text (such as project description and goals) to identifying and classifying institutions as R1 and R2, and summarizing project details into Unified Medical Language System (UMLS)-coded medical keywords. These contributions significantly reduced manual workload, allowing researchers to focus on more in-depth analysis. Our results reveal health equity insights not readily available in the original All of Us research hub. Specifically, we find a strong positive association between the use of demographic data and projects focused on health equity, while other associations such as health equity projects conducted by institutions were positive but weaker and more dependent on specific project topics.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.