Performance of a large language model (ChatGPT-3.5) for Pooled Cohort Equation estimation of atherosclerotic cardiovascular disease risk
Marafino, B. J.; Liu, V. X.
Show abstract
Despite demonstrated facility for arithmetic and other quantitative tasks, the performance of ChatGPT and other large language models for clinical risk calculation have yet to be assessed. Using synthetic patient data, this preliminary study aimed to assess the calibration, reproducibility, and potential for sociodemographic bias of ChatGPT-derived Pooled Cohort Equation (PCE) scores of atherosclerotic cardiovascular disease risk as compared to true scores. We found that ChatGPT-derived PCE scores, despite being moderately associated with the true PCE scores, displayed poor calibration with respect to true PCE scores, and exhibited instability between repeated rounds of prompting, suggesting lack of reproducibility. Moreover, ChatGPT-derived PCE scores also appeared inappropriately sensitive to contextual indicators of the sociodemographic status of the synthetic patients in this study. Further work is needed to confirm these results, and to assess performance on a wider variety of prompts as well as in other settings beyond cardiovascular disease prevention where accurate risk calculation is also vital to appropriate clinical decision-making. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/23293957v1_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@17dff17org.highwire.dtl.DTLVardef@f67a3dorg.highwire.dtl.DTLVardef@1d3665forg.highwire.dtl.DTLVardef@1e5e7e8_HPS_FORMAT_FIGEXP M_FIG C_FIG Figure. Underestimation of true PCE risk estimates (x-axis) by ChatGPT (y-axis) on synthetic patient data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 93%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 93%
- Response to Polygenic Risk: Results of the MyGeneRank Mobile Application-Based Coronary Artery Disease Study 92%
Similar papers in this journal
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 95%
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 93%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
Similar papers in this journal
Similar papers in this journal
- A scoping review of fair machine learning techniques when using real-world data 93%
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 92%
- Demonstrating the Consequences of Learning Missingness Patterns in Early Warning Systems for Preventative Health Care: A Novel Simulation and Solution 92%
Similar papers in this journal
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 91%
- Scalable information extraction from free text electronic health records using large language models 91%
- KMSubtraction: Reconstruction of unreported subgroup survival data utilizing published Kaplan-Meier survival curves 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.