Back

Bias Amplification in Intersectional Subpopulations for Clinical Phenotyping by Large Language Models

Pal, R.; Garg, H.; Patel, S.; Sethi, T.

2023-03-25 medical ethics
10.1101/2023.03.22.23287585 medRxiv
Show abstract

Large Language Models (LLMs) have demonstrated remarkable performance across diverse clinical tasks. However, there is growing concern that LLMs may amplify human bias and reduce performance quality for vulnerable subpopulations. Therefore, it is critical to investigate algorithmic underdiagnosis in clinical notes, which represent a key source of information for disease diagnosis and treatment. This study examines prevalence of bias in two datasets - smoking and obesity - for clinical phenotyping. Our results demonstrate that state-of-the-art language models selectively and consistently underdiagnosed vulnerable intersectional subpopulations such as young-aged-males for smoking and middle-aged-females for obesity. Deployment of LLMs with such biases risks skewing clinicians decision-making which may lead to inequitable access to healthcare. These findings emphasize the need for careful evaluation of LLMs in clinical practice and highlight the potential ethical implications of deploying such systems in disease diagnosis and prognosis.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.