Large-scale classification of metagenomic samples: a comparative analysis of classical machine learning techniques vs a novel brain-inspired hyperdimensional computing approach
Joshi, J. P.; Cumbo, F.; Blankenberg, D.
Show abstract
Classical machine learning techniques have revolutionized bioinformatics, enabling researchers to extract knowledge from complex biological data. However, these techniques often struggle with high-dimensional data, where the increasing number of features leads to decreased performance, also affecting models accuracy. To address this problem, we explore hyperdimensional computing (HDC), an emerging brain-inspired computational paradigm that leverages high-dimensional vectors and simple arithmetic operations to represent and manipulate complex patterns, as an alternative approach in the context of supervised machine learning. In this work, we present a comprehensive comparative analysis of HDC against established machine learning techniques across a range of classification tasks. As a representative use case, we focus on classifying heterogeneous metagenomic samples based on their quantitative microbial profiles, using publicly available microbiome datasets. Our results demonstrate that HDC achieves comparable, and in some cases, superior classification accuracy to classical methods. Furthermore, our findings highlight the potential of HDC for improved computational efficiency, particularly when dealing with large-scale datasets, suggesting the HDC-based classifier as a promising tool for bioinformatics research, particularly in areas characterized by high-dimensional data. We also offer a Galaxy powered toolset to analyze your own datasets and generate reproducible workflows and adopt these methods in your own research with ease. Our investigation into the application of a HDC-based supervised machine learning technique for classifying microbial profiles in metagenomic samples yielded promising results, demonstrating the potential of this novel computational paradigm to complement and, in some cases, surpass the performances of well established machine learning techniques. ImportanceThe growing complexity and dimensionality of biological data require more efficient and scalable machine learning approaches. HDC offers a novel alternative to conventional methods, showing resilience to high-dimensionality while maintaining competitive accuracy. This study demonstrates the effectiveness of HDC in classifying metagenomic samples based on their microbial composition. Our results suggest that HDC not only matches, but sometimes exceeds the performance of well-established methods. We make this approach accessible to the broader bioinformatics community with an open-source tool fully integrated into the Galaxy platform, facilitating its adoption and reproducibility, with the aim of integrating HDC into mainstream biological data analysis pipelines, especially for complex, high-dimensional tasks in microbiome research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Learning, Visualizing and Exploring 16S rRNA Structure Using an Attention-based Deep Neural Network 95%
- Assessing the Performance of Methods for Cell Clustering from Single-cell DNA Sequencing Data 95%
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 95%
Similar papers in this journal
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 98%
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 96%
- Assessing Random Forest self-reproducibility for optimal short biomarker signature discovery 96%
Similar papers in this journal
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 95%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 95%
- Stochastic LASSO for extremely high-dimensional genomic data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.