OmicSelector: automatic feature selection and deep learning modeling for omic experiments.
Stawiski, K.; Kaszkowiak, M.; Mikulski, D.; Hogendorf, P.; Durczynski, A.; Strzelczyk, J.; Chowdhury, D.; Fendler, W.
Show abstract
A crucial phase of modern biomarker discovery studies is selecting the most promising features from high-throughput screening assays. Here, we present the OmicSelector - Docker-based web application and R package that facilitates the analysis of such experiments. OmicSelector provides a consistent and overfitting-resilient pipeline that integrates 94 feature selection approaches based on 25 distinct variable selection methods. It identifies and then ranks the best feature sets using 11 modeling techniques with hyperparameter optimization in hold-out or cross-validation. OmicSelector provides classification performance metrics for proposed feature sets, allowing researchers to choose the overfitting-resistant biomarker set with the highest diagnostic potential. Finally, it performs GPU-accelerated development, validation, and implementation of deep learning feedforward neural networks (up to 3 hidden layers, with or without autoencoders) on selected signatures. The application performs an extensive grid search of hyperparameters, including balancing and preprocessing of next-generation sequencing (e.g. RNA-seq, miRNA-seq) oraz qPCR data. The pipeline is applicable for determining candidate circulating or tissue miRNAs, gene expression data and methylomic, metabolomic or proteomic analyses. As a case study, we use OmicSelector to develop a diagnostic test for pancreatic and biliary tract cancer based on serum small RNA next-generation sequencing (miRNA-seq) data. The tool is open-source and available at https://biostat.umed.pl/OmicSelector/
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Differential Expression Analysis with InMoose, the Integrated Multi-Omic Open-Source Environment in Python 94%
- driveR: A Novel Method for Prioritizing Cancer Driver Genes Using Somatic Genomics Data 94%
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 93%
Similar papers in this journal
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 94%
- PiDeel: Pathway-informed deep learning model for survivalanalysis and pathological classification of gliomas 93%
- CuBlock: A cross-platform normalization method for gene-expression microarrays 93%
Similar papers in this journal
- scaLR: a low-resource deep neural network-based platform for single cell analysis and biomarker discovery 94%
- Gene regulatory network integration with multi-omics data enhances survival predictions in cancer 94%
- Assessing Random Forest self-reproducibility for optimal short biomarker signature discovery 94%
Similar papers in this journal
- OmicsFootPrint: a framework to integrate and interpret multi-omics data using circular images and deep neural networks 96%
- Assessing the impact of transcriptomics data analysis pipelines on downstream functional enrichment results 95%
- miEAA 2.0: Integrating multi-species microRNA enrichment analysis and workflow management systems 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.