Are Different Populations Fairly Represented in Single-Cell Omic Atlases?
Yang, C.; Saravanan, K.; Saharan, A.; Huang, K.-l.
Show abstract
Single-cell Omic atlases are transforming biology and medicine, yet their demographic representativeness has not been systematically evaluated. We conducted a secondary analysis of >13,500 samples from the Human Cell Atlas (HCA), Human Tumor Atlas Network (HTAN), and PsychAD Consortium. Benchmarking against global and US general and disease-prevalence data, we found a striking, pervasive European over-representation and underrepresentation of Asian and Latino individuals. Nearly 70% of HCA samples lacked ancestry annotation, HTAN tumors were 69% European, and PsychAD showed more balanced non-European representation but markedly few Asians. Sex distributions in these atlases also skewed compared to published disease prevalence. These disparities highlight that current single-cell resources risk embedding inequities into AI foundational models, biomarker discovery, and therapeutic development.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Molecular Signatures of Resilience to Alzheimer's Disease in Neocortical Layer 4 Neurons 95%
- Whole-genome sequencing of 1,171 elderly admixed individuals from the largest Latin American metropolis (Sao Paulo, Brazil) 95%
- A human single-cell atlas of the Substantia nigra reveals novel cell-specific pathways associated with the genetic risk of Parkinson's disease and neuropsychiatric disorders. 94%
Similar papers in this journal
- Genotyping and population structure of the China Kadoorie Biobank 95%
- ABCA7-dependent Neuropeptide-Y signalling is a resilience mechanism required for synaptic integrity in Alzheimer's disease 94%
- Proteome-wide Mendelian randomization in global biobank meta-analysis reveals multi-ancestry drug targets for common diseases 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.