omicML: An Integrative Bioinformatics and Machine Learning Framework for Transcriptomic Biomarker Identification
Debnath, J. P.; Hossen, K.; Khandaker, M. S.; Majid, S.; Islam, M. M.; Arefin, S.; Chondrow Dev, P.; Sarker, S.; Hossain, T.
Show abstract
IntroductionTranscriptomic biomarker discovery has been a challenge due to variation in datasets and platforms, complexity in statistical and computational methods, integration of multiple programming languages, and intricacy of ML workflow to evaluate biomarkers. Standard workflows necessitate several stages (quality control, normalization, differential expression), typically executed in R or Python, resulting in bottlenecks for non-experts. Existing platforms have alleviated certain challenges by offering graphical interfaces for data loading, normalization, differential gene expression analysis, and functional analysis; nevertheless, they typically do not incorporate integrated machine learning procedures for biomarker selection. MethodIn this regard, we present omicML, an intuitive graphical user interface (GUI) that combines transcriptomic data analysis with machine learning (ML)-based classification via integrating R and Python packages/libraries. It supports both RNA-Seq and microarray data, automating preprocessing (data import, quality control, and normalization) and differential expression analysis. The tool annotates differentially expressed genes (DEGs) with descriptions, gene ontology, and pathway information and incorporates comparative analysis. Our extensive ML pipeline enables both supervised and unsupervised learning, integrates various datasets based on candidate gene signatures, standardizes and eliminates less significant features, benchmarks multiple ML classifiers with robust performance metrics (e.g., AUROC, AUPRC), assesses feature importance, develops single-gene and multi-gene predictive models, and systematically finalizes the biomarker algorithm. All functionalities are available in omicML, hence reducing the barrier for biologists without computational proficiency. ResultIn a case study, omicML identified a six-gene diagnostic model that distinguishes Mpox (monkeypox virus) infections from those caused by other viruses, including SARS-CoV-2, HIV, Ebola, and varicella-zoster. These results illustrate omicMLs capacity to discern clinically relevant biomarkers from complex transcriptome data. ConclusionThrough the unified system, omicML (https://omicml.org), integrating data preprocessing, differential gene expression analysis, annotation, heatmap analysis, dataset integration, batch effect correction, machine learning approach, and functional analysis can diminish technical barriers and accelerates the conversion of expression data into diagnostic insights for clinicians and bench scientists.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- VIGA: an one-stop tool for eukaryotic Virus Identification and Genome Assembly from next-generation-sequencing data 94%
- SPCS: A Spatial and Pattern Combined Smoothing Method of Spatial Transcriptomic Expression 94%
- Phylogeny-aware linear B-cell epitope predictor detects candidate targets for specific immune responses to Monkeypox virus 94%
Similar papers in this journal
- Enrichment analysis on regulatory subspaces: a novel direction for the superior description of cellular responses to SARS-CoV-2 95%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 94%
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 93%
Similar papers in this journal
- HIHISIV: a database of gene expression in HIV and SIV host immune response 95%
- ViralQuest: A user-friendly interactive pipeline for viral-sequences analysis and curation 94%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 93%
Similar papers in this journal
- Machine learning using intrinsic genomic signatures for rapid classification of novel pathogens: COVID-19 case study 95%
- Short k-mer Abundance Profiles Yield Robust Machine Learning Features and Accurate Classifiers for RNA Viruses 94%
- High throughput SARS-CoV-2 variant analysis using molecular barcodes coupled with Next Generation Sequencing 94%
Similar papers in this journal
- COVID-19: Viral-host interactome analyzed by network based-approach model to study pathogenesis of SARS-CoV-2 infection. 92%
- Integration of Machine Learning to Identify Diagnostic Genes in Leukocytes for Acute Myocardial Infarction Patients 92%
- RepCOOL: Computational Drug Repositioning Via Integrating Heterogeneous Biological Networks 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.