Integrating Biological and Machine Learning Models for Rainbow Trout Growth: Balancing Accuracy and Interpretability
Fulton, L.; Lyu, P.
Show abstract
Invasive species management demands predictive models that balance accuracy with ecological interpretability. Traditional approaches often fail to capture complex environmental interactions. We evaluated hybrid frameworks integrating biological and machine learning models for rainbow trout (Oncorhynchus mykiss) growth in the Lower Colorado River. Using ten years of tag-recapture data and environmental covariates, we assessed traditional and Bayesian von Bertalanffy (VBGM) and Gompertz models, Random Forests, XGBoost, LightGBM, Support Vector Regression, Neural Networks, and ensemble approaches through comprehensive probabilistic comparisons. Our results reveal substantial improvements from incorporating environmental context and advanced modeling. Top methods achieved 70 to 80 percent error reductions compared to baseline models, equivalent to 45 to 70 mm improvements or 20 to 32 percent of mean fish length. A stacked ensemble of XGBoost and the VBGM achieved optimal performance (RMSE = 15.96 mm, R2 = 0.966), demonstrating complete stochastic dominance across the posterior. Gradient boosting models formed a strong second tier, with LightGBM (9 dominances) and XGBoost (8 dominances) leading this group. Bayesian Model Averaging achieved similar accuracy while explicitly quantifying uncertainty. Even traditional mechanistic models improved markedly, up to 80 percent, when enhanced with covariates and Bayesian estimation. Feature importance analysis identified (on average) initial length, time at large, and weight at release as key predictors. The stacked ensemble dominated baseline models in over 99 percent of posterior samples, confirming its robustness. These findings establish hybrid ensemble frameworks as powerful tools for ecological forecasting, uniting predictive performance with mechanistic insight critical to conservation decision-making. The methodology provides a generalizable template for ecological systems where both accuracy and interpretability are essential. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=97 SRC="FIGDIR/small/686633v1_ufig1.gif" ALT="Figure 1"> View larger version (16K): org.highwire.dtl.DTLVardef@12c22c4org.highwire.dtl.DTLVardef@9e7796org.highwire.dtl.DTLVardef@1bd59cborg.highwire.dtl.DTLVardef@524a67_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assessing predictive accuracy of species abundance models in dynamic systems 95%
- Analysing biodiversity observation data collected in continuous time: Should we use discrete- or continuous-time occupancy models? 94%
- Dynamic Generalised Additive Models (DGAM) for forecasting discrete ecological time series 94%
Similar papers in this journal
- A machine learning method for estimating the probability of presence using presence-background data 94%
- gauseR: Simple methods for fitting Lotka-Volterra models describing Gause's "Struggle for Existence" 93%
- Modeling the demography of species providing extended parental care: A capture-recapture multievent model with a case study on Polar Bears (Ursus maritimus) 92%
Similar papers in this journal
Similar papers in this journal
- Identification and performance of environmentally-driven recruitment relationships in state space assessment models 94%
- A statistical censoring approach accounts for hook competition in abundance indices from longline surveys 94%
- Evaluating the consequences of common assumptions in run reconstructions on Pacific-salmon biological status assessments 93%
Similar papers in this journal
- Choosing priors in Bayesian ecological models by simulating from the prior predictive distribution 93%
- Confronting population models with experimental microcosm data: from trajectory matching to state-space models 92%
- The dependence of forecasts on sampling frequency as a guide to optimizing monitoring in community ecology 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.