NHANES-GPT: Large Language Models (LLMs) and the Future of Biostatistics
Titus, A. J.
Show abstract
BackgroundLarge Language Models (LLMs) like ChatGPT have significant potential in biomedicine and health, particularly in biostatistics, where they can lower barriers to complex data analysis for novices and experts alike. However, concerns regarding data accuracy and model-generated hallucinations necessitate strategies for independent verification. ObjectiveThis study, using NHANES data as a representative case study, demonstrates how ChatGPT can assist clinicians, students, and trained biostatisticians in conducting analyses and illustrates a method to independently verify the information provided by ChatGPT, addressing concerns about data accuracy. MethodsThe study employed ChatGPT to guide the analysis of obesity and diabetes trends in the NHANES dataset from 2005-2006 to 2017-2018. The process included data preparation, logistic regression modeling, and iterative refinement of analyses with confounding variables. Verification of ChatGPTs recommendations was conducted through direct statistical data analysis and cross-referencing with established statistical methodologies. ResultsChatGPT effectively guided the statistical analysis process, simplifying the interpretation of NHANES data. Initial models indicated increasing trends in obesity and diabetes prevalence in the U.S.. Adjusted models, controlling for confounders such as age, gender, and socioeconomic status, provided nuanced insights, confirming the general trends but also highlighting the influence of these factors. ConclusionsChatGPT can facilitate biostatistical analyses in healthcare research, making statistical methods more accessible. The study also underscores the importance of independent verification mechanisms to ensure the accuracy of LLM-assisted analyses. This approach can be pivotal in harnessing the potential of LLMs while maintaining rigorous standards of data accuracy and reliability in biomedical research.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- %svy_logistic_regression: A generic SAS(R) macro for simple and multiple logistic regression and creating quality publication-ready tables using survey or non-survey data 95%
- Demographic and socioeconomic determinants of access to care: A subgroup disparity analysis using new equity-focused measurements 95%
- Magnitude, pattern and correlates of multimorbidity among patients attending chronic outpatient medical care in Bahir Dar, northwest Ethiopia: the application of latent class analysis model 94%
Similar papers in this journal
- Machine learning-based equations for improved body composition estimation in Indian adults 93%
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 92%
- How can digital citizen science approaches improve ethical smartphone use surveillance among youth: traditional surveys versus ecological momentary assessments 91%
Similar papers in this journal
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 92%
- Illustrating Potential Effects of Alternate Control Populations on Real-World Evidence-based Statistical Analyses 92%
- Characterization and Racial Stratification of Social Determinants of Health for Individuals with Type 2 Diabetes as Recorded in Electronic Health Records: Implications for Artificial Intelligence Development 92%
Similar papers in this journal
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 93%
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 91%
- Machine Learning Directed Interventions Associate with Decreased Hospitalization Rates in Hemodialysis Patients 90%
Similar papers in this journal
- Computational Strategies in Nutrigenetics: Constructing a Reference Dataset of Nutrition-Associated Genetic Polymorphisms 93%
- Creating an Ignorance-Base: Exploring Known Unknowns in the Scientific Literature 90%
- Impact of COVID-19 on the outreach strategy of cancer social service agencies in Singapore: A pre-post analysis with Facebook data 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.