Interpretable and predictive models to harness the life science data revolution
Jahner, J. P.; Buerkle, C. A.; Gannon, D. G.; Grames, E. M.; McFarlane, S. E.; Siefert, A.; Bell, K. L.; DeLeo, V. L.; Forister, M. L.; Harrison, J. G.; Laughlin, D. C.; Patterson, A. C.; Powers, B. F.; Werner, C. M.; Oleksy, I. A.
Show abstract
The proliferation of high-dimensional data in ecology and evolutionary biology raises the promise of statistical and machine learning models that are highly predictive and interpretable. However, high-dimensional data are commonly burdened with an inherent trade-off: in-sample prediction of outcomes will improve as additional variables are included in the model, but this may come at the cost of poor predictive accuracy and limited generalizability for future or unsampled observations (out-of-sample prediction). To confront this problem of overfitting, sparse models can focus on key variables by correctly placing low weight on unimportant variables. We competed nine methods to quantify their performance in variable selection and prediction using simulated data with different sample sizes, numbers of variables, and strengths of effects. Overfitting was typical for many methods and simulation scenarios. Despite this, in-sample and out-of-sample prediction converged on the true predictive target for simulations with more observations, larger causal effects, and fewer variables. Accurate variable selection to support process-based understanding will be unattainable for many realistic sampling schemes in ecology and evolution. We use our analyses to characterize data attributes for which statistical learning is possible, and illustrate how some sparse methods can achieve predictive accuracy while mitigating and learning the extent of overfitting.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A multi-view model for relative and absolute microbial abundances 93%
- An Interpretable Bayesian Clustering Approach with Feature Selection for Analyzing Spatially Resolved Transcriptomics Data 93%
- A Bayesian Approach to Restricted Latent Class Models for Scientifically-Structured Clustering of Multivariate Binary Outcomes 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.