Back

A universal single-cell transcriptomics atlas of human lung decodes multiple pulmonary diseases

Wu, F.; Cai, W.; Tang, H.; Zheng, S.; Zhang, H.; Chen, Y.; Han, Y.; Zhou, D.; Wang, R.; Ye, M.; You, R.; Chen, A.; Li, J.; Zhang, X.; Li, W.

2024-12-23 bioinformatics
10.1101/2024.12.17.628654 bioRxiv
Show abstract

Human lung is a complex organ susceptible to various diseases. Single-cell transcriptomic studies provide rich data to targeting specific research questions. Here, we present uniLUNG, the largest lung transcriptomic cell atlas, comprising over 10 million cells across 20 disease states and healthy controls. We ensembled a universal hierarchical annotation framework and conducted a full benchmarking of data integration to define a standardized nomenclature and marker genes for lung cell types. Using uniLUNG, we identified Lym-monocyte and T-like B cells, new cell types in specific lung diseases, confirming their existence by comparing with external single-cell atlases. Additionally, we discovered the NSCLC-like SCLC subpopulation, a transitional malignant cell population associated with the transition from NSCLC to SCLC, which was validated and further characterized in spatial dimensions, revealing its complex role in tumour progression. Overall, uniLUNG represents a comprehensive range of human lung cell diversity, providing valuable data resources and a reliable foundation for lung single-cell research. HIGHLIGHTSO_LIThe largest scRNA atlas for human lung covers 10 million cells from 20 lung states. C_LIO_LIA four-level universal cell annotation framework encompasses 120 lung cell types. C_LIO_LIComprehensive benchmarking on 18 strategies guides data integration. C_LIO_LISpecific distribution of Lym-monocytes and T-like B cells in specific lung diseases. C_LIO_LIThe NSCLC-like SCLC subpopulation in transitional events of malignant cells from NSCLC to SCLC. C_LI

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.