Back

ButterflyVI: enabling high-throughput variant interpretation and biomarker discovery with functional genomics

Sesia, D.; Hatos, A.; De Nittis, P.; Ciriello, G.

2026-01-22 cancer biology
10.64898/2026.01.20.700339 bioRxiv
Show abstract

Characterizing the functional and therapeutic relevance of cancer mutations is a primary goal and a major challenge in precision oncology. Whereas predictive approaches exist, they often lack functional validation, which cannot scale to large numbers of cancer-associated variants. Here, we integrated statistical learning with experimental evidence from high-throughput functional screenings to annotate >20000 unique variants. We defined a variant as functional if it altered the effect of gene loss, and classified four possible outcomes: oncogene and tumor suppressor dependencies, mutation tolerance, and bypass-of-essentiality. Up-to-60% of variants annotated as functional were previously considered of unknown significance. Paradoxically, bypass-of-essentiality was common among loss-of-function (LoF) variants at several tumor suppressors, including VHL, ARID1A, and RBM10. In these cases, loss of the wild-type genes was deleterious, independently of the tissue of origin, but not when they already harbored recurrent LoF variants, suggesting loss of these tumor suppressors provides a context-specific advantage. Using our annotations, we discovered somatic variants that increased sensitivity to loss of therapeutically targetable genes, representing new candidate biomarkers. Among these, we validated RPL5 LoF mutations as a biomarker of response to selective MDM2 inhibitors. A dedicated web portal (butterflyvi.unil.ch) enables exploration of all variant annotations and candidate biomarkers. This study highlights the potential and need to expand large-scale functional screenings to empower variant interpretation in the clinic.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.