The Effects of Nonlinear Signal on Expression-Based Prediction Performance
Heil, B. J.; Crawford, J.; Greene, C. S.
10.1101/2022.06.22.497194 bioRxivShow abstract
Those building predictive models from transcriptomic data are faced with two conflicting perspectives. The first, based on the inherent high dimensionality of biological systems, supposes that complex non-linear models such as neural networks will better match complex biological systems. The second, imagining that complex systems will still be well predicted by simple dividing lines prefers linear models that are easier to interpret. We compare multi-layer neural networks and logistic regression across multiple prediction tasks on GTEx and Recount3 datasets and find evidence in favor of both possibilities. We verified the presence of non-linear signal for transcriptomic prediction tasks by removing the predictive linear signal with Limma, and showed the removal ablated the performance of linear methods but not non-linear ones. However, we also found that the presence of non-linear signal was not necessarily sufficient for neural networks to outperform logistic regression. Our results demonstrate that while multi-layer neural networks may be useful for making predictions from gene expression data, including a linear baseline model is critical because while biological systems are highdimensional, effective dividing lines for predictive models may not be.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Iterative point set registration for aligning scRNA-seq data 95%
- Randomized Spatial PCA (RASP): a computationally efficient method for dimensionality reduction of high-resolution spatial transcriptomics data 95%
- A Bayesian method to infer copy number clones from single-cell RNA and ATAC sequencing 95%
Similar papers in this journal
Similar papers in this journal
- multiDGD: A versatile deep generative model for multi-omics data 95%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 95%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 95%
Similar papers in this journal
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 96%
- An efficient not-only-linear correlation coefficient based on machine learning 95%
- Belayer: Modeling discrete and continuous spatial variation in gene expression from spatially resolved transcriptomics 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.