Back

Race, Ethnicity and Their Implication on Bias in Large Language Models

Hu, S.; Li, R.; Gao, Y.

2026-01-05 health informatics
10.64898/2026.01.04.26343415 medRxiv
Show abstract

Large language models (LLMs) increasingly operate in high-stakes settings including healthcare and medicine, where demographic attributes such as race and ethnicity may be explicitly stated or implicitly inferred from text. However, existing studies primarily document outcome-level disparities, offering limited insight into internal mechanisms underlying these effects. We present a mechanistic study of how race and ethnicity are represented and operationalized within LLMs. Using two publicly available datasets spanning toxicity-related generation and clinical narrative understanding tasks, we analyze three open-source models with a re-producible interpretability pipeline combining probing, neuron-level attribution, and targeted intervention. We find that demographic information is distributed across internal units with substantial cross-model variation. Although some units encode sensitive or stereotype-related associations from pretraining, identical demographic cues can induce qualitatively different behaviors. Interventions suppressing such neurons reduce bias but leave substantial residual effects, suggesting behavioral rather than representational change and motivating more systematic mitigation.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

1
Nature Communications
4913 papers in training set
Top 29%
6.3%
2
Scientific Reports
3102 papers in training set
Top 19%
6.3%
3
Nature
575 papers in training set
Top 5%
4.8%
4
Nature Machine Intelligence
61 papers in training set
Top 0.6%
4.8%
5
PNAS Nexus
147 papers in training set
Top 0.1%
4.8%
6
Philosophical Transactions of the Royal Society B
51 papers in training set
Top 0.7%
4.8%
7
Science Advances
1098 papers in training set
Top 4%
3.9%
8
Nature Medicine
117 papers in training set
Top 0.8%
3.6%
9
eLife
5422 papers in training set
Top 26%
3.6%
10
Science
429 papers in training set
Top 9%
3.6%
11
Proceedings of the National Academy of Sciences
2130 papers in training set
Top 20%
3.6%
50% of probability mass above
12
Nature Biomedical Engineering
42 papers in training set
Top 0.4%
2.7%
13
npj Digital Medicine
97 papers in training set
Top 2%
2.6%
14
Cell Systems
167 papers in training set
Top 5%
2.6%
15
Nature Human Behaviour
85 papers in training set
Top 1%
2.6%
16
PLOS Computational Biology
1633 papers in training set
Top 13%
2.4%
17
GENETICS
189 papers in training set
Top 0.4%
2.3%
18
Science Translational Medicine
111 papers in training set
Top 2%
2.1%
19
Cell Reports Medicine
140 papers in training set
Top 3%
1.9%
20
Med
38 papers in training set
Top 0.3%
1.7%
21
Neuron
282 papers in training set
Top 6%
1.7%
22
iScience
1063 papers in training set
Top 15%
1.7%
23
PLOS ONE
4510 papers in training set
Top 57%
1.5%
24
Communications Psychology
20 papers in training set
Top 0.1%
1.3%
25
Journal of the American Medical Informatics Association
61 papers in training set
Top 1%
1.3%
26
Annals of Internal Medicine
27 papers in training set
Top 0.5%
1.3%
27
Communications Biology
886 papers in training set
Top 14%
1.2%
28
Nature Neuroscience
216 papers in training set
Top 5%
1.2%
29
Advanced Science
249 papers in training set
Top 14%
1.2%
30
Communications Medicine
85 papers in training set
Top 0.9%
0.8%