Prediction of Alzheimer's Disease from Single Cell Transcriptomics Using Deep Learning
Srivastava, A.; Dhall, A.; Patiyal, S.; Arora, A.; Jarwal, A.; Raghava, G. P. S.
Show abstract
Alzheimers disease (AD) is a progressive neurological disorder characterized by brain cell death, brain atrophy, and cognitive decline. Early diagnosis of AD remains a significant challenge in effectively managing this debilitating disease. In this study, we aimed to harness the potential of single-cell transcriptomics data from 12 Alzheimers patients and 9 normal controls (NC) to develop a predictive model for identifying AD patients. The dataset comprised gene expression profiles of 33,538 genes across 169,469 cells, with 90,713 cells belonging to AD patients and 78,783 cells belonging to NC individuals. Employing machine learning and deep learning techniques, we developed prediction models. Initially, we performed data processing to identify genes expressed in most cells. These genes were then ranked based on their ability to classify AD and NC groups. Subsequently, two sets of genes, consisting of 35 and 100 genes, respectively, were used to develop machine learning-based models. Although these models demonstrated high performance on the training dataset, their performance on the validation/independent dataset was notably poor, indicating potential overoptimization. To address this challenge, we developed a deep learning method utilizing dropout regularization technique. Our deep learning approach achieved an AUC of 0.75 and 0.84 on the validation dataset using the sets of 35 and 100 genes, respectively. Furthermore, we conducted gene ontology enrichment analysis on the selected genes to elucidate their biological roles and gain insights into the underlying mechanisms of Alzheimers disease. While this study presents a prototype method for predicting AD using single-cell genomics data, it is important to note that the limited size of the dataset represents a major limitation. To facilitate the scientific community, we have created a website to provide with code and service. It is freely available at https://webs.iiitd.edu.in/raghava/alzscpred. Key PointsO_LIPredictive Model for Alzheimers Disease Using Single Cell Transcriptomics Data C_LIO_LIOveroptimization of models trained on single-cell genomics data. C_LIO_LIApplication of dropout regularization technique of ANN for reducing overoptimization C_LIO_LIRanking of genes based on their ability to predict patients Alzheimers Disease C_LIO_LIStandalone software package for predicting Alzheimers Disease C_LI Authors BiographyO_LIAman Srivastava is pursuing M. Tech. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAnjali Dhall is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LISumeet Patiyal is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAkanksha Arora is currently working as Ph.D. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIAkanksha Jarwal is pursuing M. Tech. in Computational Biology from Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIGajendra P. S. Raghava is currently working as Professor and Head of Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of functionally connected multi-omic biomarkers for Alzheimer’s Disease using modularity-constrained Lasso 97%
- Random forest model for feature-based Alzheimer's disease conversion prediction from early mild cognitive impairment subjects 96%
- c-Triadem: A constrained, explainable deep learning model to identify novel biomarkers in Alzheimer’s disease 96%
Similar papers in this journal
- Network-based identification of genetic factors in Ageing, lifestyle and Type 2 Diabetes that Influence in the progression of Alzheimer’s disease 91%
- An Inexpensive Smartphone-Based Device and Predictive Models for Rapid, Non-Invasive, and Point-of-Care Monitoring of Ocular and Cardiovascular Complications Related to Diabetes 91%
- Finding Consensus miRNAs Silencing KLF1 Expression as A Promising Therapeutic Option of Sickle Cell Anemia 91%
Similar papers in this journal
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 93%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 93%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 92%
Similar papers in this journal
- Common molecular signatures between coronavirus infection and Alzheimer's disease reveal targets for drug development 96%
- Quantitative longitudinal predictions of Alzheimer's disease by multi-modal predictive learning 95%
- Targeted Serum Metabolomic Profiling and Machine Learning Approach in Alzheimer’s Disease using the Alzheimer’s Disease Diagnostics Clinical Study (ADDIA) Cohort 92%
Similar papers in this journal
- Artificial intelligence-driven meta-analysis of brain gene expression data identifies novel gene candidates in Alzheimer’s Disease 93%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 92%
- Multi-task deep autoencoder to predict Alzheimer’s disease progression using temporal DNA methylation data in peripheral blood 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.