Back

Development and validation of a reliable DNA copy-number-based machine learning algorithm (CopyClust) for breast cancer integrative cluster classification

Young, C. C.; Eason, K.; Manzano Garcia, R.; Moulange, R.; Mukherjee, S.; Chin, S.-F. C.; Caldas, C.; Rueda, O. M.

2023-11-22 genomics
10.1101/2023.11.21.568129 bioRxiv
Show abstract

The Integrative Clusters (IntClusts) provide a framework for the classification of breast cancer tumors into 10 distinct genomic subtypes based on DNA copy number and gene expression. Current classifiers achieve only low accuracy without gene expression data, warranting the development of new approaches to copy-number-only-based IntClust classification. A novel XGBoost-driven classification algorithm, CopyClust, was trained using genomic features from METABRIC and validated on TCGA achieving a nine-percentage point or greater improvement in overall IntClust subtype classification accuracy.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.