Back

Characterization and classification of fine-resolution soil profile for precision agriculture using random forest and self-organizing map

Elias, A. A.; Sharma, M.; Goel, S.

2024-04-03 plant biology
10.1101/2024.04.02.587707 bioRxiv
Show abstract

The availability of high throughput soil profile information is an important component in precision agriculture to perform efficient soil management for sustainable production. We collected 14 soil physiochemical features from Nagpur, Pune, and Haveri, representing target environments of safflower cultivation and also from our experiment station at Delhi, at fine resolution and created graphical maps to depict the variability. Additionally, we evaluated the predictive ability of two statistical learning models, random forest (RF) and self-organizing maps (SOM) against multinomial regression models for correctly classifying the soil profile. Clustering was performed around the medoids produced from the dissimilarity matrices of these models using partitioning around medoids (PAM) model. The robustness, versatility, and predictive ability of models in correctly classifying the soil profile to clusters were then tested using cross-validation which was repeated 100 times. This study was performed using training data with proportionate size varying from 60 to 95%, and increasing the unit area of observation up to nine times (or decreasing the total number of observations up to a ninth). RF model was found to be the best performing with average prediction accuracy above 85% in all settings which reached close to 100% in some settings. The predictive ability of all the models was maintained even when only the most influencing six variables were used for classification. The optimal training population size for prediction was found to be 70 - 80%. Based on our study, it is recommended to i) collect fine resolution edaphic features from a marginal farm before crop season, ii) use RF or SOM model to identify the most influencing features distinguishing the soil samples iii) expand the area of sample collection, find values for the most influencing features, and use RF model to correctly predict the class to which the new set of the soil belongs to. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=151 SRC="FIGDIR/small/587707v1_ufig1.gif" ALT="Figure 1"> View larger version (47K): org.highwire.dtl.DTLVardef@2f404forg.highwire.dtl.DTLVardef@2707c2org.highwire.dtl.DTLVardef@6e6418org.highwire.dtl.DTLVardef@16da481_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.