AlzGenPred: A CatBoost based method using network features to classify the Alzheimers Disease associated genes from the high throughput sequencing data
Shukla, R.; Singh, T. R.
Show abstract
Background and ObjectiveAD is a progressive neurodegenerative disorder characterized by memory loss. Due to the advancement in next-generation sequencing technologies, an enormous amount of AD-associated genomics data is available. However, the information about the involvement of these genes in AD association is still a research topic because all these algorithms are based on statistical techniques. Therefore, AlzGenPred is developed to identify the AD-associated genes from a large set of data. MethodsTo develop the AlzGenPred, we have compiled a benchmark dataset consisting of 1086 AD and non-AD genes and used them as positive and negative datasets. We have generated several features including the fused features and evaluated them through machine learning methods. Then hyperparameter tuning approach was also applied and the final model was selected. The proposed method was validated by using the AlzGene and transcriptomics datasets and proposed as a standalone tool. ResultsTotal 13504 features belonging to eight different encoding schemes of these sequences were generated and evaluated by using 16 ML algorithms. It reveals that network-based features can classify AD genes while sequence-based features are not able to classify them. Then we generated 24 different fused features (6020 D) using sequence-based features and fed them into a two-step lightGBM-based recursive feature selection method. It increased up to 5-7% accuracy. After that selected eight fused features with CKSAAP were used for the hyperparameter tuning. They showed <70% accuracy. Therefore, network-based features were used to generate the CatBoost-based ML method called AlzGenPred with 96.55% accuracy and 98.99% AUROC. The developed method is tested on the AlzGene dataset where it showed 96.43% accuracy. Then the model is validated using the transcriptomics dataset also. ConclusionThe validation of AlzGenPred using the AlzGene dataset and transcriptomics dataset obtained from Human, mouse, and ES-derived neural cells revealed that it can classify the omics data and can sort the AD-associated genes. These predicted genes can be directly used in the wet lab for further testing which will reduce labor cost and time expenses. The AlzGenPred is developed as a standalone package and is available for users at https://www.bioinfoindia.org/alzgenpred/ and https://github.com/shuklarohit815/AlzGenPred.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 96%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 95%
- A method for predicting linear and conformational B-cell epitopes in an antigen from its primary sequence 94%
Similar papers in this journal
- Identification of functionally connected multi-omic biomarkers for Alzheimer’s Disease using modularity-constrained Lasso 96%
- c-Triadem: A constrained, explainable deep learning model to identify novel biomarkers in Alzheimer’s disease 95%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
Similar papers in this journal
- Extensive In Silico Analysis of the Functional and Structural Consequences of SNPs in Human ARX Gene 92%
- Genome-wide identification and prediction of SARS-CoV-2 mutations show an abundance of variants: Integrated study of bioinformatics and deep neural learning. 92%
- A Multiple Peptides Vaccine against nCOVID-19 Designed from the Nucleocapsid phosphoprotein (N) and Spike Glycoprotein (S) via the Immunoinformatics Approach 92%
Similar papers in this journal
- AITeQ: A machine learning framework for Alzheimer's prediction using a distinctive 5-gene signature 98%
- SPCS: A Spatial and Pattern Combined Smoothing Method of Spatial Transcriptomic Expression 95%
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 94%
Similar papers in this journal
- iMDA-BN: Identification of miRNA-Disease Associations based on the Biological Network and Graph Embedding Algorithm 94%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 94%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.