Back

Evaluating reasoning LLMs' potential to perpetuate racial and gender disease stereotypes in healthcare

Docking, J. J.; Li, L. X.; Menz, B. D.; Bacchi, S.; Hopkins, A. M.; Sorich, M. J.

2025-08-07 health informatics
10.1101/2025.08.05.25333007 medRxiv
Show abstract

This evaluation of 36,000 clinical vignettes found that next-generation reasoning large language models, o3-mini and DeepSeek-R1, frequently perpetuate racial and gender stereotypes for common medical conditions, indicating that advancements in reasoning do not inherently improve representational fairness.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.