Approaches to dimensionality reduction for ultra-high dimensional models
Kotlarz, K.; Slomian, D.; Szyda, J.
Show abstract
The rapid advancement of high-throughput sequencing technologies has revolutionised genomic research by providing access to large amounts of genomic data. However, the most important disadvantage of using Whole Genome Sequencing (WGS) data is its statistical nature, the so-called p>>n problem. This study aimed to compare three approaches of feature selection allowing for circumventing the p>>n problem, among which one is a novel modification of Supervised Rank Aggregation (SRA). The use of the three methods was demonstrated in the classification of 1,825 individuals representing the 1000 Bull Genomes Project to 5 breeds, based on 11,915,233 SNP genotypes from WGS. In the first step, we applied three feature (i.e. SNP) selection methods: the mechanistic approach (SNP tagging) and two approaches considering biological and statistical contexts by fitting a multiclass logistic regression model followed by either 1-dimensional clustering (1D-SRA) or multi-dimensional feature clustering (MD-SRA) that was originally proposed in this study. Next, we perform the classification based on a Deep Learning architecture composed of Convolutional Neural Networks. The classification quality of the test data set was expressed by macro F1-Score. The SNPs selected by SNP tagging yielded the least satisfactory results (86.87%). Still, this approach offered rapid computing times by focussing only on pairwise LD between SNPs and disregarding the effects of SNP on classification. 1D-SRA was less suitable for ultra-high-dimensional applications due to computational, memory and storage limitations, however, the SNP set selected by this approach provided the best classification quality (96.81%). MD-SRA provided a very good balance between classification quality (95.12%) and computational efficiency (17x lower analysis time and 14x lower data storage), outperforming other methods. Moreover, unlike SNP tagging, both SRA-based approaches are universal and not limited to feature selection for genomic data. Our work addresses the urgent need for computational techniques that are both effective and efficient in the analysis and interpretation of large-scale genomic datasets. We offer a model suitable for the classification of ultra-high-dimensional data that implements fusing feature selection and deep learning techniques.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BWGS: a R package for genomic selection and its application to a wheat breeding programme. 95%
- Cluster analysis on high dimensional RNA-seq data with applications to cancer research- An evaluation study 94%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 94%
Similar papers in this journal
- MLcps: Machine Learning Cumulative Performance Score for classification problems 94%
- parSMURF, a High Performance Computing tool for the genome-wide detection of pathogenic variants 93%
- SnpHub: an easy-to-set-up web server framework for exploring large-scale genomic variation data in the post-genomic era with applications in wheat 93%
Similar papers in this journal
- Coupling Day Length Data and Genomic Prediction tools for Predicting Time-Related Traits under Complex Scenarios 94%
- Use of a graph neural network to the weighted gene co-expression network analysis of Korean native cattle 93%
- Novel AI-powered computational method using tensor decomposition for identification of common optimal bin sizes when integrating multiple Hi-C datasets 93%
Similar papers in this journal
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 96%
- Assessing Random Forest self-reproducibility for optimal short biomarker signature discovery 94%
- Advances in multi-trait genomic prediction approaches: Classification, comparative analysis, and perspectives 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.