Back

CanBART: A generative foundation model of cancer molecular alterations for synthetic patient generation and genomic profile completion

Chebanov, D.; Morris, Q. D.

2025-11-06 oncology
10.1101/2025.11.04.25339512 medRxiv
Show abstract

Real-world genomic cohorts remain limited in size, unevenly profiled, and especially sparse for rare cancers. Many sequencing panels differ in gene coverage, leaving key molecular features unobserved and complicating cohort comparison, biomarker discovery, and clinical-trial enrollment. CanBART, a generative foundation model trained on 144,000 tumor profiles, learns tokenized representations of somatic alterations and supports genomic completion, tumor classification, and synthetic-patient generation via a masked-language architecture. We simulated a panel reduction scenario by limiting each MSK-IMPACT profile to the smaller set of genes measured in the DFCI OncoPanel, and asked the model to reconstruct the missing genes. In this setting, CanBART recovered over one-third of the held-out alterations with high confidence. Using sampling strategies adapted from NLP, we also generated biologically coherent "plausible patients" that expanded rare cancer cohorts, and training classifiers on these synthetic profiles improved accuracy for two-thirds of tumor types, particularly those with only 20-500 samples. By completing missing genomic features and generating "plausible" synthetic patients, CanBART provides a scalable tool for rare cancer research and cross-panel harmonization.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.