Back

Single and multi-analyte deep learning-based analysis framework for class prediction in biological images

Krishnan, N. M.; Kumar, S.; Kumar, U.; Panda, B.

2022-10-17 plant biology
10.1101/2022.10.13.512074 bioRxiv
Show abstract

Measurement of biological analytes, characterizing flavor in fruits, is a cumbersome, expensive and time-consuming process. Fruits with higher concentration of analytes have greater commercial or nutritional values. Here, we tested a deep learning-based framework with fruit images to predict the class (sweet or sour and high or low) of analytes using images from two types of trees in a single and multi-analyte mode. We used fruit images from kinnow (n = 3,451), an edible hybrid mandarin and neem (n = 1,045), a tree with agrochemical and pharmaceutical properties. We measured sweetness in kinnows and five secondary metabolites in neem fruits (azadirachtin or A, deacetyl-salannin or D, salannin or S, nimbin or N and nimbolide or E) using a refractometer and high-performance liquid chromatography, respectively. We trained the models for 300 epochs, before and after hyper-parameters evolution, using 300 generations with 50 epochs/generation, estimated the best models and evaluated their performance on 10% of independent images. The validation F1score and test accuracies were 0.79 and 0.77, and 82.55% and 60.8%, respectively for kinnow and neem A analyte. A multi-analyte model enhanced the neem A models prediction to high class when the D:N:Ss combined class predictions were high:low:high and to low class when D:Ns combined class predictions were low:high respectively. The test accuracy increased further to ~70% with a 10-fold cross-validation error of 0.257 across ten randomly split train:validation:test sets proving the potential of a multi-analyte model to enhance the prediction accuracy, especially when the numbers of images are limiting.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.