Back

DPCGS: a computational framework for linking GWAS to single-cell transcriptomics in complex traits and diseases

Liu, C.; Shen, B.; Li, J.; Zhu, R.; Yang, P.; Wu, B.; Xuan, Y.; Yang, S.; Yuan, B.; Yang, N.; Ma, L.; Liu, Q.; Dai, S.; Zhang, Y.

2026-07-20 bioinformatics
10.64898/2026.07.14.738331 bioRxiv
Show abstract

Complex traits and diseases arise from the interplay between genetic variation and cellular heterogeneity, making it essential to understand how genetic risk manifests at the cellular level. However, connecting genome-wide association studies (GWAS) to specific cell populations remains challenging due to cellular complexity and the prevalence of noncoding variants. Here, we present DPCGS, a computational framework that systematically integrates GWAS summary statistics with single-cell RNA-sequencing (scRNA-seq) data to identify trait-associated cell populations, genes, and regulatory programs. DPCGS is based on the principle that GWAS-prioritized genes should exhibit elevated expression in relevant cells compared with matched controls. Benchmark analyses with simulated datasets showed that DPCGS consistently outperforms existing methods, achieving higher accuracy and sensitivity in detecting trait-relevant cells. Applications to diverse scRNA-seq datasets further validated its robustness, revealing oligodendrocytes and astrocytes as key subpopulations in Alzheimers disease and macrophages and B cells in asthma. These analyses also highlighted potential molecular regulators, including CD74, FOS, FLI1, and AP-1 transcription factors. Together, these findings establish DPCGS as a versatile framework for dissecting the cellular and molecular basis of complex traits and diseases, with broad implications for biomarker discovery and therapeutic development.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.