Back

Beyond Histology: A Unified Transcriptomic Atlas Defines Lung Cancer Biologic States and Subtypes

Arora, S.; Suresh, L.; Thirmanne, H. N.; Jensen, M.; Glatzer, G.; Fatherree, J.; Konnick, E.; Levine, K.; Brooks, A. N.; Houghton, A. M.; Pritchard, C.; MacPherson, D.; Berger, A.; Holland, E. C.

2026-03-18 bioinformatics
10.64898/2026.03.16.712177 bioRxiv
Show abstract

Lung cancer encompasses multiple histological entities with substantial molecular heterogeneity that remain incompletely resolved at population scale. Here, we constructed a unified reference landscape of lung cancer by analyzing raw RNA sequencing data from 1,558 tumors spanning adenocarcinoma (n=753), squamous cell carcinoma (n=540), small cell lung cancer (n=150), and unclassified non-small cell lung cancer (n=80). Following batch correction, samples were embedded using PaCMAP to generate a continuous molecular atlas annotated with clinical and biological metadata. Rather than segregating strictly by histology, tumors organized along conserved transcriptional axes defined by tumor-intrinsic proliferative or metabolic programs and immune-infiltrated states. Consensus clustering resolved nine robust molecular clusters, including a female non-smoker-enriched adenocarcinoma subgroup, a neuroendocrine-like adenocarcinoma marked by ASCL1 activation, immune-associated regions, and bifurcation of both small cell and squamous carcinomas into biologically distinct states. Spatially-restricted expression of clinically actionable targets revealed state-specific vulnerabilities. Projection of patient tumors and patient-derived xenografts onto the atlas demonstrated preservation of transcriptional identity and enabled quantitative assessment of model fidelity. This unified framework redefines lung cancer as a structured continuum of transcriptional states with translational relevance. One Sentence SummaryA landscape built using only transcriptomic analysis for lung cancer reveals novel insights about subtype-specific biology.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.