Sociodemographic Bias in Large Language Model Clinical Trial Screening
Soffer, S.; Omar, M.; efros, o.; Apakama, D. U.; Mudrik, A.; Freeman, R.; Nadkarni, G.; Klang, E.
Show abstract
BackgroundLarge language models (LLMs) are increasingly used in randomized clinical trial (RCT) screening, but their potential for sociodemographic bias remains unclear. ObjectiveTo determine whether LLM-based trial screening judgments vary with patient sociodemographic characteristics when clinical details and eligibility criteria are held constant. Design, Setting, and ParticipantsCross-sectional evaluation of Phase II-III RCT protocols from ClinicalTrials.gov (U.S. adult populations; 2023-2024). For each protocol, we created 15 physician-validated clinical vignettes rendered in 34 versions: one control (no identifiers) and 33 identity variants spanning gender, race/ethnicity, socioeconomic status, homelessness, unemployment, and sexual orientation. ExposuresIdentity labels applied to otherwise identical vignettes, evaluated by nine contemporary LLMs. Main Outcomes and MeasuresPrimary eligibility domain score (1-5 Likert scale) comparing identity variants versus control. Secondary: adherence, resources, risk-benefit, and trust/attitude domains. Mixed-effects models estimated adjusted mean differences with multiplicity-corrected P values; differences <.10 considered trivial. ResultsOf 69 protocols, 58 met inclusion criteria. Analysis of 5,324,400 model evaluations showed eligibility judgments were largely stable: most identity-related differences fell within {+/-}0.05 (transgender woman -.008 [95% CI -.04 to .02]; White male .036 [.01 to .07]). Only homelessness exceeded the trivial threshold (-.121 [-.15 to -.09], P<.001). Secondary domains revealed socioeconomic gradients, particularly for adherence (homeless -.595, P<.001) and resources (homeless -.715, P<.001), with smaller trust/attitude effects and negligible risk-benefit differences. Conclusions and RelevanceBias in LLM-assisted trial screening is conditional. Within fixed criteria, models reason consistently; outside them, they echo the inequities of their data. Responsible deployment in clinical research depends on preserving that boundary so that automation strengthens fairness in trial access rather than inheriting distortion.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Results dissemination from clinical trials conducted at German university medical centres was delayed and incomplete 92%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 92%
- Results reporting for clinical trials led by medical universities and university hospitals in the Nordic countries was often missing or delayed 91%
Similar papers in this journal
- Effect of the Friendship Bench intervention on antiretroviral therapy outcomes and mental health symptoms in rural Zimbabwe: A cluster randomized trial 91%
- Effect of Montelukast vs Placebo on Time to Sustained Recovery in Outpatients with COVID-19: The ACTIV-6 Randomized Clinical Trial 91%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 91%
Similar papers in this journal
- New Model, Old Risks? Sociodemographic Bias and Adversarial Hallucinations Vulnerability in GPT-5 91%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 91%
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 91%
Similar papers in this journal
- How Informative Were Early SARS-CoV-2 Treatment and Prevention Trials? A longitudinal cohort analysis of trials registered on clinicaltrials.gov 93%
- Hydroxychloroquine/Chloroquine for the Treatment of Hospitalized Patients with COVID-19: An Individual Participant Data Meta-Analysis 92%
- STIMULATE-ICP: A pragmatic, multi-centre, cluster randomised trial of an integrated care pathway with a nested, Phase III, open label, adaptive platform randomised drug trial in individuals with Long COVID: a structured protocol 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.