Back

Qombucha: Reconstructing unobserved progenitor methylation profiles reveals distinct developmental programs in glioblastoma

Li, X. C.; Lalchungnunga, H.; Hari, A.; Liu, Y.; Singh, O.; Wu, Z.; Abdullaev, Z.; Mount, S. M.; Aldape, K. D.; Ruppin, E.; Schaffer, A. A.; Sahinalp, S. C.

2026-08-12 bioinformatics
10.64898/2026.08.06.743400 bioRxiv
Show abstract

Glioblastoma (GBM) is a highly aggressive brain cancer characterized by substantial intratumoral heterogeneity. Previous research demonstrates that GBM may have complex cell origins. To elucidate the interplay between brain development and GBM progression, we introduce Qombucha (Quadratic prOgraMming Based tUmor deConvolution with cell HierArchy), a computational framework that uses DNA methylation data to infer tumor cell-type composition and profiles of unobserved progenitor cells. Unprecedentedly, Qombucha incorporates a developmental cell hierarchy that models mature brain cell types and their progenitors. Applied to a large TCGA GBM dataset spanning the RTK I, RTK II, and MES TYP subtypes, Qombucha identifies a distinct cell type composition profile for each subtype and recapitulates known biological patterns, including elevated microglia infiltration in MES TYP tumors and increased RTK I stemness. It also identifies subtype-specific developmental programs and shows that higher progenitor-cell abundance is associated with poorer survival. Qombucha-imputed cell fractions map methylation profiles of tumor samples to a compact, 11-dimensional latent space; in an independent NCI GBM cohort, this compact representation improves subtype clustering and enables accurate subtype classification, achieving performance comparable to stateof-the-art models based on full methylation profiles with much higher dimensionality. These results suggest that Qombucha-imputed tumor cellular composition captures the core biological axes along which GBM subtypes diverge.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.