Back

Sparse input neural networks to differentiate 32 primary cancer types based on somatic point mutations

Dikaios, N.

2020-05-15 bioinformatics
10.1101/2020.05.13.092916 bioRxiv
Show abstract

This paper aims to differentiate cancer types from primary tumour samples based on somatic point mutations (SPM). Primary cancer site identification is necessary to perform site-specific and potentially targeted treatment. Current methods like histopathology/lab-tests cannot accurately determine cancers origin, which results in empirical patient treatment and poor survival rates. The availability of large deoxyribonucleic-acid sequencing datasets has allowed scientists to examine the ability of SPM to classify primary cancer sites. These datasets are highly sparse since most genes will not be mutated, have low signal-to-noise ratio and are imbalanced since rare cancers have less samples. To overcome these limitations a sparse-input neural network (spinn) is suggested that projects the input data in a lower dimensional space, where the more informative genes are used for learning. To train and evaluate spinn, an extensive dataset was collected from the cancer genome atlas containing 7624 samples spanning 32 cancer types. Different sampling strategies were performed to balance the dataset but have not benefited the classifiers performance except for removing Tomek-links. This is probably due to high amount of class overlapping. Spinn consistently outperformed algorithms like extreme gradient-boosting, deep neural networks and support-vector-machines, achieving an accuracy up to 73% on independent testing data.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.