Visualizing the Costs and Benefits of Correcting P-Values for Multiple Hypothesis Testing in Omics Data
Shuken, S. R.; McNerney, M. W.
Show abstract
The multiple hypothesis testing problem is inherent in high-throughput quantitative genomic, transcriptomic, proteomic, and other "omic" screens. The correction of p-values for multiple testing is a critical element of quantitative omic data analysis, yet many researchers are unfamiliar with the sensitivity costs and false discovery rate (FDR) benefits of p-value correction. We developed models of quantitative omic experiments, modeled the costs and benefits of p-value correction, and visualized the results with color-coded volcano plots. We developed an R Shiny web application for further exploration of these models which we call the Simulator of P-value Multiple Hypothesis Correction (SIMPLYCORRECT). We modeled experiments in which no analytes were truly differential between the control and test group (all null hypotheses true), all analytes were differential, or a mixture of differential and non-differential analytes were present. We corrected p-values using the Benjamini-Hochberg (BH), Bonferroni, and permutation FDR methods and compared the costs and benefits of each. By manipulating variables in the models, we demonstrated that increasing sample size or decreasing variability can reduce or eliminate the sensitivity cost of p-value correction and that permutation FDR correction can yield more hits than BH-adjusted and even unadjusted p-values in strongly differential data. SIMPLYCORRECT can serve as a tool in education and research to show how p-value adjustment and various parameters affect the results of quantitative omics experiments.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Biological Function Assignment Across Taxonomic Levels in Mass-Spectrometry-Based Metaproteomics via a Modified Expectation Maximization Algorithm 94%
- Bimodal peptide collision cross section distribution reflects two stable conformations in the gas phase 94%
- Quality control for the target decoy approach for peptide identification 94%
Similar papers in this journal
- SmartPeak automates targeted and quantitative metabolomics data processing 95%
- Concepts and Software Package for Efficient Quality Control in Targeted Metabolomics Studies - MeTaQuaC 94%
- mzrtsim: Raw Data Simulation for Reproducible Gas/Liquid Chromatography Mass Spectrometry Based Non-targeted Metabolomics Data Analysis 93%
Similar papers in this journal
Similar papers in this journal
- Pathway analysis in metabolomics: pitfalls and best practice for the use of over-representation analysis 94%
- Common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline 93%
- PathIntegrate: Multivariate modelling approaches for pathway-based multi-omics data integration 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.