Analysis of variance when both input and output sets are high-dimensional
de los Campos, G. A.; Torsten, P.; Gonzalez-Raymundez, A.; Simianer, H.; Mias, G.; Vazquez, A. I.
Show abstract
MotivationModern genomic data sets often involve multiple data-layers (e.g., DNA-sequence, gene expression), each of which itself can be high-dimensional. The biological processes underlying these data-layers can lead to intricate multivariate association patterns. ResultsWe propose and evaluate two methods for analysis variance when both input and output sets are high-dimensional. Our approach uses random effects models to estimate the proportion of variance of vectors in the linear span of the output set that can be explained by regression on the input set. We consider a method based on orthogonal basis (Eigen-ANOVA) and one that uses random vectors (Monte Carlo ANOVA, MC-ANOVA) in the linear span of the output set. We used simulations to assess the bias and variance of each of the methods, and to compare it with that of the Partial Least Squares (PLS)-an approach commonly used in multivariate-high-dimensional regressions. The MC-ANOVA method gave nearly unbiased estimates in all the simulation scenarios considered. Estimates produced by Eigen-ANOVA and PLS had noticeable biases. Finally, we demonstrate insight that can be obtained with the of MC-ANOVA and Eigen-ANOVA by applying these two methods to the study of multi-locus linkage disequilibrium in chicken genomes and to the assessment of inter-dependencies between gene expression, methylation and copy-number-variants in data from breast cancer tumors. AvailabilityThe Supplementary data includes an R-implementation of each of the proposed methods as well as the scripts used in simulations and in the real-data analyses. Contactgustavoc@msu.edu Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adding gene transcripts into genomic prediction improves accuracy and reveals sampling time dependence 95%
- Interpretable Artificial Neural Networks incorporating Bayesian Alphabet Models for Genome-wide Prediction and Association Studies 95%
- A deep learning framework for characterization of genotype data 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 94%
- Accelerated matrix-vector multiplications for matrices involving genotype covariates with applications in genomic prediction 93%
- Machine Learning Approaches Identify Genes Containing Spatial Information from Single-Cell Transcriptomics Data. 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.