Asymmetric sociodemographic disparity in evidence-grounded clinical AI
Jia, E.; Omar, M.; Barash, Y.; Brook, O. R.; Ahmed, M.; Kruskal, J. B.; Gorenshtein, A.; Klang, E.
Show abstract
AI-assisted clinical care may compound, rather than correct, existing health inequities. We applied Omar and colleagues' validated four-domain emergency-medicine benchmark to OpenEvidence (OE), a literature-grounded clinical LLM used by tens of thousands of US physicians daily, across 100 emergency-department cases and 20 sociodemographic labels. OE was consistent on the codified clinical decisions, triage, workup, and treatment, but diverged sharply on mental-health screening, where it flagged many historically marginalized groups between three and ten times more often than demographically unmarked cases. Cases labeled as unhoused received recommendations in 78 to 87 percent of responses (versus a 9 percent no-identifier-control rate); cases labeled as transgender in 22 to 24 percent; and Black transgender women specifically in 47 percent. A pre- registered audit of 193 free-text rationales localized the differential to the inner layer of the response, in the structure and tone of the rationale rather than the recommendation itself. Literature grounding may redistribute sociodemographic disparity in clinical AI rather than remove it. As clinical LLMs move toward agentic deployment, equity audits should examine how evidence is applied to each patient, not only whether citations are present.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multi-ancestry study of the genetics of problematic alcohol use in >1 million individuals 90%
- Identification of 64 new risk loci for major depression, refinement of the genetic architecture and risk prediction of recurrence and comorbidities 90%
- Polygenic risk scores lack prognostic value for adults with severe mental illness 90%
Similar papers in this journal
- New Model, Old Risks? Sociodemographic Bias and Adversarial Hallucinations Vulnerability in GPT-5 93%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 91%
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 90%
Similar papers in this journal
- Clinical features and burden of post-acute sequelae of SARS-CoV-2 infection in children and adolescents: an exploratory EHR-based cohort study from the RECOVER program 86%
- Longitudinal associations between weight indices, cognition, and mental health from childhood to early adolescence 86%
- Impacts of school closures on physical and mental health of children and young people: a systematic review 86%
Similar papers in this journal
- Correlations between human mating partners: a comprehensive meta-analysis of 22 traits and raw data analysis of 133 traits in the UK Biobank 91%
- Indecision and recency-weighted evidence integration in non-clinical and clinical settings 91%
- Correction for participation bias in the UK Biobank reveals non-negligible impact on genetic associations and downstream analyses 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.