Back

Data leakage and measurement error inflate the apparent predictability of overyielding from plant traits

Kopp, E. B.; Koenig, N.; Vonmetz, L.; Wuest, S. E.; Niklaus, P. A.

2026-01-15 ecology
10.64898/2026.01.15.699684 bioRxiv
Show abstract

Predicting plant mixture overyielding from functional traits is central to understanding how biodiversity influences ecosystem functioning. Combining empirical data from a large mixture experiment (764 mixtures) and two trait-measurement experiments containing 90 soybean genotypes, with complementary simulations where trait-function relationships and noise levels were defined a priori, we show that apparent model performance depends critically on how predictive models are validated. When training and testing data share genotypes, predictive ability is strongly inflated because shared monoculture and trait measurements create data leakage. Measurement errors further propagate through these shared components, inducing spurious correlations and amplifying noise. When validation uses completely independent genotypes, the predictive power of both linear and machine-learning models declines sharply, revealing limited but genuine predictability. These results show how data structure and measurement error can produce misleading model performance and underscore the need for rigorous validation to achieve robust ecological prediction.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.