Evaluating reasoning LLMs' potential to perpetuate racial and gender disease stereotypes in healthcare
Docking, J. J.; Li, L. X.; Menz, B. D.; Bacchi, S.; Hopkins, A. M.; Sorich, M. J.
Show abstract
This evaluation of 36,000 clinical vignettes found that next-generation reasoning large language models, o3-mini and DeepSeek-R1, frequently perpetuate racial and gender stereotypes for common medical conditions, indicating that advancements in reasoning do not inherently improve representational fairness.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 93%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 92%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 93%
Similar papers in this journal
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 94%
- COVID-19 outcomes, risk factors and associations by race: a comprehensive analysis using electronic health records data in Michigan Medicine 93%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 92%
Similar papers in this journal
- Development and Application of Pharmacological Statin-Associated Muscle Symptoms Phenotyping Algorithms Using Structured and Unstructured Electronic Health Records Data 92%
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 91%
- Algorithmic Individual Fairness and Healthcare: A Scoping Review 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.