Back

Scaling and Generalization of Discrete Diffusion Models for Tumor Phylogenies

Sabata, S.; Schwartz, R.

2026-03-26 bioinformatics
10.64898/2026.03.23.713822 bioRxiv
Show abstract

Tumor phylogenies -- rooted trees encoding clonal ancestry and mutation acquisition -- are central to understanding cancer evolution, yet generating realistic phylogenies remains challenging. We investigate whether discrete graph diffusion can learn the structural constraints of tumor phylogenies directly from data. Working with approximately 12,500 synthetic phylogenies across twelve evolutionary regimes, we train graph transformer models that denoise typed graphs through a learned reverse diffusion process. Scaling experiments reveal a non-monotonic capacity-performance relationship: a mid-scale model achieves high structural validity and close distributional match to held-out data, while a deeper model fails under fixed optimization hyperparameters. Low-data cross-regime experiments show that diverse training produces more transferable representations than single-regime specialization. These results establish that phylogenetic structural constraints can be learned implicitly through unconditional discrete diffusion, suggesting a viable path toward generative models of tumor evolution.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.