Identifying and Characterizing Bias at Scale in Clinical Notes Using Large Language Models
Apakama, D. U.; Nguyen, K.-A.-N.; Hyppolite, D.; Soffer, S.; Mudrik, A.; Ling, E.; Moses, A.; Temnycky, I.; Glasser, A.; Anderson, R.; Parchure, P.; Woullard, E.; Edalati, M.; Chan, L.; Kronk, C.; Freeman, R.; Kia, A.; Timsina, P.; Levin, M.; Khera, R.; Patricia Kovatch, P.; Charney, A. W.; Carr, B. G.; Richardson, L. D.; Horowitz, C. R.; Klang, E.; Nadkarni, G.
Show abstract
ImportanceDiscriminatory language in clinical documentation impacts patient care and reinforces systemic biases. Scalable tools to detect and mitigate this are needed. ObjectiveDetermine utility of a frontier large language model (GPT-4) in identifying and categorizing biased language and evaluate its suggestions for debiasing. DesignCross-sectional study analyzing emergency department (ED) notes from the Mount Sinai Health System (MSHS) and discharge notes from MIMIC-IV. SettingMSHS, a large urban healthcare system, and MIMIC-IV, a public dataset. ParticipantsWe randomly selected 50,000 ED medical and nursing notes from 230,967 MSHS 2023 adult ED visiting patients, and 500 randomly selected discharge notes from 145,915 patients in MIMIC-IV database. One note was selected for each unique patient. Main Outcomes and MeasuresPrimary measure was accuracy of detection and categorization (discrediting, stigmatizing/labeling, judgmental, and stereotyping) of bias compared to human review. Secondary measures were proportion of patients with any bias, differences in the prevalence of bias across demographic and socioeconomic subgroups, and provider ratings of effectiveness of GPT-4s debiasing language. ResultsBias was detected in 6.5% of MSHS and 7.4% of MIMIC-IV notes. Compared to manual review, GPT-4 had sensitivity of 95%, specificity of 86%, positive predictive value of 84% and negative predictive value of 96% for bias detection. Stigmatizing/labeling (3.4%), judgmental (3.2%), and discrediting (4.0%) biases were most prevalent. There was higher bias in Black patients (8.3%), transgender individuals (15.7% for trans-female, 16.7% for trans-male), and undomiciled individuals (27%). Patients with non-commercial insurance, particularly Medicaid, also had higher bias (8.9%). Higher bias was also seen in health-related characteristics like frequent healthcare utilization (21% for >100 visits) and substance use disorders (32.2%). Physician-authored notes showed higher bias than nursing notes (9.4% vs. 4.2%, p < 0.001). GPT-4s suggested revisions were rated highly effective by physicians, with an average improvement score of 9.6/10 in reducing bias. Conclusions and RelevanceA frontier LLM effectively identified biased language, without further training, showing utility as a scalable fairness tool. High bias prevalence linked to certain patient characteristics underscores the need for targeted interventions. Integrating AI to facilitate unbiased documentation could significantly impact clinical practice and health outcomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Impact of Primary Care Team Configuration on Access and Quality of Care 93%
- Association of patients’ past misdiagnosis experiences with trust in their current physician: the TRUMP 2 -Net study 92%
- An Evaluation of the Vulnerable Physician Workforce in the United States During the Coronavirus Disease-19 Pandemic 91%
Similar papers in this journal
- Enhancing Research Data Infrastructure to Address the Opioid Epidemic: The Opioid Overdose Network (02-Net) 94%
- Characterization and Racial Stratification of Social Determinants of Health for Individuals with Type 2 Diabetes as Recorded in Electronic Health Records: Implications for Artificial Intelligence Development 92%
- Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes 92%
Similar papers in this journal
- The Impact of the “Muslim Ban” Executive Order on Healthcare Utilization in Minneapolis-St. Paul, Minnesota 93%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 92%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 92%
Similar papers in this journal
- Predictors of adverse outcome in patients with suspected COVID-19 managed in a ‘virtual hospital’ setting: a cohort study 93%
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 92%
- Cohort Profile: A national prospective cohort study of SARS-CoV-2 pandemic outcomes in the U.S. - The CHASING COVID Cohort Study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.