scEPS integrates genetic and single-cell disease atlas data to provide granular mechanistic insights into complex human diseases
Zou, L.; Whitley, O.; Tseng, H.-W.; Simopoulos, C.; Chang, D.; Zhang, R.; Stockwell, A.; Gong, W.; Fletez-Brant, K.; Lucas, T.; Kenigsberg, E.; Missarova, A.; Garfield, D.; Yaspan, B.; McCarthy, M.; Mahajan, A.; Kaminker, J.; Shi, H.
Show abstract
Integrating GWAS and single-cell data holds great potential for prioritizing causal disease biology at cellular resolution. Recent integrative approaches typically assess the enrichment of disease genetic signals in cell types or individual cells, without directly modeling disease phenotypes. We develop a new method, single-cell Expression exPlainability Statistics (scEPS), for identifying disease-associated cell neighborhoods, by explicitly testing whether the expression of GWAS-prioritized genes explains more variance in a disease than randomly selected, mean-expression-matched control genes. Crucially, when applied to PRSs of healthy donors, scEPS captures the genetic covariance between gene expression and diseases, mitigating the effect of reverse causation and prioritizing cell populations mediating the effects of GWAS genes. We applied scEPS to clinical diagnoses and PRSs of 4 neurological and 4 respiratory disorders, integrating brain and lung cell atlas data, respectively, with respective GWAS summary statistics data. scEPS recapitulated known and uncovered novel disease-associated cell populations, identifying 1.77x (s.e. 1.21) and 5.13x (s.e. 3.08) more significant associations than a CNA-based approach and scDRS, respectively. Furthermore, scEPS detected different cell populations, contrasting clinical diagnoses vs. their PRSs, revealing distinct biology for the active/symptomatic vs. preclinical/asymptomatic states of the disease. Finally, we observed limited concordance across methods using distinct definitions of disease association, underscoring the need to integrate complementary insights for holistic understanding of disease biology.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 96%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 95%
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 95%
Similar papers in this journal
- Unsupervised representation learning improves genomic discovery and risk prediction for respiratory and circulatory functions and diseases 95%
- Atlas of genetic effects in human microglia transcriptome across brain regions, aging and disease pathologies 94%
- Dissecting tumor transcriptional heterogeneity from single-cell RNA-seq data by generalized binary covariance decomposition 94%
Similar papers in this journal
- scGRNom: a computational pipeline of integrative multi-omics analyses for predicting cell-type disease genes and regulatory networks 94%
- Leveraging genomic diversity for discovery in an EHR-linked biobank: the UCLA ATLAS Community Health Initiative 94%
- Diagnostic Evidence GAuge of Single cells (DEGAS): A flexible deep-transfer learning framework for prioritizing cells in relation to disease 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.