Back

Machine learning predictions improve identification of real-world cancer driver mutations

Tran, T. N.; Fong, C.; Pichotta, K.; Luthra, A.; Shen, R.; Chen, Y.; Waters, M.; Kim, S.; Berger, M. F.; Riely, G.; Ladanyi, M.; Chakravarty, D.; Schultz, N.; Jee, J.

2024-04-01 genomics
10.1101/2024.03.31.587410 bioRxiv
Show abstract

Characterizing and validating which mutations influence development of cancer is challenging. Machine learning has delivered significant advances in protein structure prediction, but its utility for identifying cancer drivers is less explored. We evaluated multiple computational methods for identifying cancer driver alterations. For identifying known drivers, methods incorporating protein structure or functional genomic data outperformed methods trained only on evolutionary data. We further validated VUSs annotated as pathogenic by testing their association with overall survival in two cohorts of patients with non-small cell lung cancer (N=7,965 and 977). "Pathogenic" VUSs in KEAP1 and SMARCA4 identified by several methods were associated with worse survival, unlike "benign" VUSs. "Pathogenic" VUSs exhibited mutual exclusivity with known oncogenic alterations at the pathway level, further suggesting biological validity. Despite training primarily on germline, rather than somatic, mutation data, computational predictions contribute to a more comprehensive understanding of tumor genetics as validated by real-world data.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.