ChatGPT Exhibits Gender and Racial Biases in Acute Coronary Syndrome Management
Zhang, A.; Yuksekgonul, M.; Guild, J.; Zou, J.; Wu, J.
Show abstract
Recent breakthroughs in large language models (LLMs) have led to their rapid dissemination and widespread use. One early application has been to medicine, where LLMs have been investigated to streamline clinical workflows and facilitate clinical analysis and decision-making. However, a leading barrier to the deployment of Artificial Intelligence (AI) and in particular LLMs has been concern for embedded gender and racial biases. Here, we evaluate whether a leading LLM, ChatGPT 3.5, exhibits gender and racial bias in clinical management of acute coronary syndrome (ACS). We find that specifying patients as female, African American, or Hispanic resulted in a decrease in guideline recommended medical management, diagnosis, and symptom management of ACS. Most notably, the largest disparities were seen in the recommendation of coronary angiography or stress testing for the diagnosis and further intervention of ACS and recommendation of high intensity statins. These disparities correlate with biases that have been observed clinically and have been implicated in the differential gender and racial morbidity and mortality outcomes of ACS and coronary artery disease. Furthermore, we find that the largest disparities are seen during unstable angina, where fewer explicit clinical guidelines exist. Finally, we find that through asking ChatGPT 3.5 to explain its reasoning prior to providing an answer, we are able to improve clinical accuracy and mitigate instances of gender and racial biases. This is among the first studies to demonstrate that the gender and racial biases that LLMs exhibit do in fact affect clinical management. Additionally, we demonstrate that existing strategies that improve LLM performance not only improve LLM performance in clinical management, but can also be used to mitigate gender and racial biases.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-sensitivity cardiac troponin on presentation to rule out myocardial infarction: a stepped-wedge cluster randomised controlled trial 93%
- Standardized Data Elements for Patients with Acute Pulmonary Embolism: A Consensus Report from the Pulmonary Embolism Research Collaborative 92%
- Screening for Atrial Fibrillation in Older Adults at Primary Care Visits: the VITAL-AF Randomized Controlled Trial 91%
Similar papers in this journal
- Artificial intelligence of arterial Doppler waveforms to predict major adverse outcomes among patients evaluated for peripheral artery disease 94%
- Association of Cardiologist Clinic Visits with Cardiovascular Primary Prevention Outcomes Among People with HIV from Underrepresented Racial and Ethnic Groups in the Southern United States 93%
- Long-term Outcomes of Peripheral Artery Disease In Veterans: Analysis of the PEripheral ARtery Disease Long-term Survival Study (PEARLS) 93%
Similar papers in this journal
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 94%
- Response to Polygenic Risk: Results of the MyGeneRank Mobile Application-Based Coronary Artery Disease Study 92%
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 92%
Similar papers in this journal
- The Impact of COVID-19 Pandemic on Cardiology Services 94%
- Comparison of troponin and natriuretic peptides in takotsubo syndrome and acute coronary syndrome: a meta-analysis 92%
- Impact of COVID-19 pandemic on rates of congenital heart disease procedures among children: Prospective cohort analyses of 26,270 procedures in 17,860 children using CVD-COVID-UK consortium record linkage data 92%
Similar papers in this journal
- International Evaluation Of An Artificial Intelligence-Powered Ecg Model Detecting Occlusion Myocardial Infarction 95%
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 93%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.