The Fallacy and Bias of Averages on Vegetation Indices based Plant Phenotyping
Lee, C.-H.; Tseng, C.-C.; Liu, L.-y. D.; Chen, T.-W.
Show abstract
BackgroundVegetation indices (VIs) from remote sensing are widely used for non-destructive plant phenotyping, often averaged across plots or image regions to represent each plot. However, according to Jensens inequality, which is known as the "fallacy of the average", it can bias estimates when nonlinear relationships exist between VIs and target traits. To examine this issue, we systematically assessed the severity of this bias and tested a correction method. VI values were simulated using six beta distributions with varying shapes and skewness, and with normalized difference vegetation index (NDVI) images from a paddy rice experiment to evaluate bias under real conditions. Nonlinear link functions (concave, convex, logistic) with different noise levels were applied to model VI-trait relationships. ResultThe results showed that averaging under nonlinear relationships reduced predictive performance, lowering the coefficient of determination (R2) between true and predicted traits by up to 82%. In the rice NDVI simulation, R2 was reduced by up to 58% around the tillering stage. Our correction method, which predicts traits from VI before averaging, substantially mitigated bias, improving R2 by up to 0.68 depending on noise level, VI distribution, and link function. To facilitate application, we established an interactive R Shiny website enabling users to quantify potential biases and the efficacy of corrections within this workflow based on their own research conditions ConclusionIn summary, averaging VIs without accounting for nonlinear relationships can introduce substantial bias and degrade phenotyping accuracy. This bias should be explicitly considered in phenotyping analyses, and correction methods applied when appropriate to improve reliability.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Crop loss identification at field parcel scale using satellite remote sensing and machine learning 96%
- An evaluation model for aboveground biomass Based on Hyperspectral Data from field and TM8 in Khorchin grassland, China 95%
- Suitability of resampled multispectral datasets for mapping flowering plants in the Kenyan savannah 93%
Similar papers in this journal
- High-throughput Phenotyping of Soybean Biomass: Conventional Trait Estimation and Novel Latent Feature Extraction Using UAV Remote Sensing and Deep Learning Models 96%
- Easy MPE: Extraction of quality microplot images for UAV-based high-throughput field phenotyping 95%
- SegVeg: Segmenting RGB images into green and senescent vegetation by combining deep and shallow methods 94%
Similar papers in this journal
- Developmental Normalization of Phenomics Data Generated by High Throughput Plant Phenotyping Systems 95%
- Monitoring of drought stress and transpiration rate using proximal thermal and hyperspectral imaging in an indoor automated plant phenotyping platform 94%
- Spatial and Texture Analysis of Root System Distribution with Earth Mover's Distance (STARSEED) 94%
Similar papers in this journal
- A UAV-based high-throughput phenotyping approach to assess time-series nitrogen responses and identify traits associated genetic components in maize 93%
- Quantifying Leaf Symptoms of Sorghum Charcoal Rot in Images of Field-Grown Plants 9 Using Deep Neural Networks 93%
- In-field whole plant maize architecture characterized by Latent Space Phenotyping 93%
Similar papers in this journal
- Widely Used Variants of the Farquhar-von-Caemmerer-Berry Model Can Cause Errors in Parameter Estimation 92%
- Incorporating A Dynamic Gene-Based Process Module Into A Crop Simulation Model 92%
- CRONOSOJA: a daily time-step hierarchical model predicting soybean development across maturity groups in the Southern Cone 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.