K-means Based Unsupervised Feature Selection to Prioritize Biomarkers of Different Disease Clinical Phases
Jiang, X.; Wang, W.; Xu, J.; Wang, Z.; Lin, G. N.
Show abstract
Huntingtons disease is caused by a single gene mutation, which is potentially a good model for development of biomarkers corresponding to different disease phase and clinical phenotypes. Hypothesis-driven and omics discovery approaches have not yet identified effective candidate biomarkers in HD. So, it is urgent to develop engagement and disease-phase specific biomarkers. The advanced sequencing technology makes it possible to develop data-driven methods for biomarkers discovery. Therefore, in this study, we designed k-means based unsupervised feature selection (KFS) method to prioritize biomarkers of different disease clinical phases. KFS first conducts k-means clustering on the samples with gene expression data, then it conducts feature selection based on the feature selection matrix to prioritize biomarkers of different samples. By conducting alternative iteration of clustering and feature selection to screen key genes which corresponding to the complex clinical phenotypes of different disease phases. Further gene ontology and enrichment analysis highlight potential molecular mechanisms of HD. Our experimental analyses have uncovered new disease-related genes and disease-associated pathways, which in turn have provided insight into the molecular mechanisms during the disease progression.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Regional medical inter-institutional cooperation in medical provider network constructed using patient claims data from Japan 95%
- Prediction and control of COVID-19 infection based on a hybrid intelligent model 95%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
Similar papers in this journal
- GCNCDA: A New Method for Predicting CircRNA-Disease Associations Based on Graph Convolutional Network Algorithm 97%
- A new machine learning method for cancer mutation analysis 97%
- scPADGRN: A preconditioned ADMM approach for reconstructing dynamic gene regulatory network using single-cell RNA sequencing data 96%
Similar papers in this journal
Similar papers in this journal
- BertNDA: a Model Based on Graph-Bert and Multi-scale Information Fusion for ncRNA-disease Association Prediction 96%
- scASK: A novel ensemble framework for classifying cell types based on single-cell RNA-seq data 95%
- LncDLSM: Identification of Long Non-coding RNAs with Deep Learning-based Sequence Model 95%
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 95%
- ISMI-VAE: A Deep Learning Model for Classifying Disease Cells Using Gene Expression and SNV Data 94%
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.