Experimental Design Issues Associated With Classifications Of Hyperspectral Imaging Data
Nansen, C.; Mesgaran, M.; Lee, H.
Show abstract
Hyperspectral imaging has emerged as a pivotal tool to classify plant materials (seeds, leaves, and whole plants), pharmaceutical products, food items, and many other objects. This communication addresses two issues, which appear to be over-looked or ignored in >99% of hyperspectral imaging studies: 1) the "small N, large P" problem, when number of spectral bands (explanatory variables, "P") surpasses number of observations, ("N") leading to potential model over-fitting, and 2) absence of independent validation data in performance assessments of classification models. Based on simulations of randomly generated data, we illustrate risks associated with these issues. We explore and discuss consequences of over-fitting and risks of misleadingly high accuracy that can result from having a large number of variables relative to observations. We highlight connections of these issues with radiometric repeatability (levels of stochastic noise). A method is proposed wherein a theoretical dataset is generated to mirror the structure of an actual dataset, with the classification of this theoretical dataset serving as a reference. By shedding light on important and common experimental design issues, we aim to enhance methodological rigor and transparency in classifications of hyperspectral imaging data and foster improved and effective applications across various science domains.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Open Soil Spectral Library (OSSL): Building reproducible soil calibration models through open development and community engagement 94%
- Beyond Traditional Methods: Innovative Integration of LISS IV and Sentinel 2A Imagery for Unparalleled Insight into Himalayan Ibex Habitat Suitability 93%
- Suitability of resampled multispectral datasets for mapping flowering plants in the Kenyan savannah 92%
Similar papers in this journal
- Transects, quadrats, or points? What is the best combination to get a precise estimation of a coral community? 89%
- Vocal complexity in the long calls of Bornean orangutans 87%
- Reassessing the observational evidence for nitrogen deposition impacts in acid grassland: Spatial Bayesian linear models indicate small and ambiguous effects on species richness 87%
Similar papers in this journal
- A new method for counting reproductive structures in digitized herbarium specimens using Mask R-CNN 90%
- A user manual to measure gas diffusion kinetics in plants: Pneumatron construction, operation and data analysis 90%
- Root System Architecture and Environmental Flux Analysis in Mature Crops using 3D Root Mesocosms 89%
Similar papers in this journal
- Making WAVES in Breedbase: An Integrated Spectral Data Storage and Analysis Pipeline for Plant Breeding Programs 93%
- Low-cost, handheld near-infrared spectroscopy for root dry matter content prediction in cassava 92%
- Robustness of high-throughput prediction of leaf ecophysiological traits using near infra-red spectroscopy and poro-fluorometry 92%
Similar papers in this journal
- Optimizing forest canopy structure retrieval from smartphone-based hemispherical photography 94%
- One leaf for all: Chemical traits of single leaves measured at the leaf surface using Near infrared-reflectance spectroscopy (NIRS) 93%
- imageseg: an R package for deep learning-based image segmentation 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.