Prediction of breast cancer proteins using molecular descriptors and artificial neural networks: a focus on cancer immunotherapy proteins, metastasis driver proteins, and RNA-binding proteins
Lopez-Cortes, A.; Cabrera-Andrade, A.; Vazquez-Naya, J. M.; Pazos, A.; Gonzales-Diaz, H.; Paz-y-Mino, C.; Guerrero, S.; Perez-Castillo, Y.; Tejera, E.; Munteanu, C. R.
Show abstract
BackgroundBreast cancer (BC) is a heterogeneous disease characterized by an intricate interplay between different biological aspects such as ethnicity, genomic alterations, gene expression deregulation, hormone disruption, signaling pathway alterations and environmental determinants. Due to the complexity of BC, the prediction of proteins involved in this disease is a trending topic in drug design. MethodsThis work is proposing accurate prediction classifier for BC proteins using six sets of protein sequence descriptors and 13 machine learning methods. After using a univariate feature selection for the mix of five descriptor families, the best classifier was obtained using multilayer perceptron method (artificial neural network) and 300 features. ResultsThe performance of the model is demonstrated by the area under the receiver operating characteristics (AUROC) of 0.980 {+/-} 0.0037 and accuracy of 0.936 {+/-} 0.0056 (3-fold cross-validation). Regarding the prediction of 4504 cancer-associated proteins using this model, the best ranked cancer immunotherapy proteins related to BC were RPS27, SUPT4H1, CLPSL2, POLR2K, RPL38, AKT3, CDK3, RPS20, RASL11A and UBTD1; the best ranked metastasis driver proteins related to BC were S100A9, DDA1, TXN, PRNP, RPS27, S100A14, S100A7, MAPK1, AGR3 and NDUFA13; and the best ranked RNA-binding proteins related to BC were S100A9, TXN, RPS27L, RPS27, RPS27A, RPL38, MRPL54, PPAN, RPS20 and CSRP1. ConclusionsThis powerful model predicts several BC-related proteins which should be deeply studied to find new biomarkers and better therapeutic targets. The script and the results are available as a free repository at https://github.com/muntisa/neural-networks-for-breast-cancer-proteins.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Directed Bayesian Networks established functional differences between breast cancer subtypes 95%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 94%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 94%
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 97%
- Risk assessment of cancer patients based on HLA-I alleles, neobinders and expression of cytokines 95%
- A method for predicting linear and conformational B-cell epitopes in an antigen from its primary sequence 94%
Similar papers in this journal
- Unveiling epigenetic regulatory elements associated with breast cancer development 97%
- Multi-run Concrete Autoencoder to Identify Prognostic lncRNAs for 12 Cancers 95%
- Data independent acquisition mass spectrometry (DIA-MS) analysis of FFPE rectal cancer samples offers in depth proteomics characterization of response to neoadjuvant chemoradiotherapy 94%
Similar papers in this journal
- Classification models for Invasive Ductal Carcinoma Progression, based on gene expression data-trained supervised machine learning 97%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 95%
- Machine learning prediction of antiviral-HPV protein interactions for anti-HPV pharmacotherapy 95%
Similar papers in this journal
- Analysis of Mutations in Precision Oncology using The Automated, Accurate, and User-Friendly Web Tool PredictONCO 94%
- Hybrid computational modeling highlights reverse Warburg effect in breast cancer-associated fibroblasts 94%
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.