Back

Machine learning-based analysis of genomic and transcriptomic data unveils sarcoma clusters with superlative prognostic and predictive value

Esperanca-Martins, M.; Vasques, H.; Ravasqueira, M. S.; Lemos, M. M.; Fonseca, F.; Coutinho, D.; Lopez, J. A.; Huang, R. S. P.; Dias, S.; Gallego-Paez, L.; Costa, L.; Abecasis, N.; Goncalves, E.; Fernandes, I.

2025-02-02 oncology
10.1101/2025.01.31.25321492 medRxiv
Show abstract

Soft tissue sarcomas (STS) histopathological classification system has several conceptual caveats, impacting prognostication and treatment. The clinical and molecular-based tools currently employed to estimate prognosis also have limitations. Clinically driven molecular profiling studies may cover these gaps. We performed DNA sequencing (DNAseq) and RNA sequencing (RNAseq), portraying the molecular profile of 102 samples of 3 of the most common STS subtypes. The RNAseq data was analyzed using unsupervised machine learning models, unravelling previously unknown molecular patterns and identifying 4 well-defined transcriptomic clusters. These transcriptomic clusters have a clear prognostic value, a finding that was externally validated. This transcriptomic cluster-based classifications prognostic value is superior to the prognostic accuracy of currently used clinical-based (SARCULATOR nomograms) and molecular-based (CINSARC) prognostication tools. The analysis of DNAseq data from the same cohort of samples revealed a plethora of unique and, in some cases, never documented molecular targets for precision treatment across different transcriptomic clusters.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.