Back

Single-Cell Genomics Elucidates Molecular Variations and Regulatory Mechanisms in Circulating Immune Cells

Yin, J.; Zheng, Y.; Huang, Z.; Zhou, W.; Yuan, Y.; Cai, P.; Bai, Y.; Yang, S.; Gao, Y.; Duan, S.; Wang, Y.; Zhang, W.; Zhang, X.; Wei, Y.; Xu, Z.; Huang, Y.; Liu, Y.; Wang, W.; Yang, T.; Lv, J.; Zhang, Z.; Chen, X.; Zhang, X.; Li, F.; Zhang, Y.; Zeng, G.; Wang, X.; Ma, W.; Hou, G.; Hao, S.; Liu, C.; Lai, Y.; Wang, B.; Li, Y.; Zhang, W.; Gao, P.; Xie, J.; Esteban, M. A.; Gu, Y.; Ji, J.; Qi, T.; Liu, B.; Wang, J.; Yang, J.; Xu, X.; Liu, L.; Jin, X.; Liu, C.

2025-01-27 immunology
10.1101/2025.01.26.634963 bioRxiv
Show abstract

The human peripheral blood displays diverse molecular characteristics across populations, understanding the drivers and underlying mechanisms of which remains challenging. Here, we introduce the Chinese Immune Multi-Omics Atlas (CIMA), elucidating sex-, age-, and genetic-related molecular variations by analyzing multi-omics data from 428 adults with over 10 million immune cells. CIMA generated an enhancer-driven gene regulatory network, identifying 237 high-quality regulons and revealing cell type-specific regulatory mechanisms. Additionally, 11,521 lead cis-expression quantitative trait loci (eQTLs) and 46,339 chromatin accessibility QTLs (caQTLs) were identified at cell type level. CIMA also uncovered pleiotropic associations among immune-related disease risk loci, eQTLs, and caQTLs in a cell type-specific manner. Lastly, a novel cell language model, CIMA-CLM, was developed to predict chromatin accessibility and noncoding variant effects using chromatin sequences and gene expressions. This work represents a population-scale multi-omics resource of human immune cells, providing a valuable reference for future investigation of immune-related diseases.

Published in Science (predicted rank #7) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.