Back

Odorant molecular feature mining by diverse deep neural networks for prediction of odor perception categories

Shang, L.; Liu, C.; Tang, F.; Chen, B.; Liu, L.; Hayashi, K.

2022-04-22 bioinformatics
10.1101/2022.04.20.488977 bioRxiv
Show abstract

The use of an artificial intelligence (AI)-based prediction model of the structure-odor relationship (SOR) has shown great potential in the replacement of human panelists in gas chromatography-olfactometry (GCO). However, the Al-based GCO encounters issues such as poor accuracy, generalization, and practicality, owning to the insufficient feature extraction of odorant molecular structure. The purpose of this study is to the prediction of odor perception categories based on the odorant structure feature extraction by diverse deep neural networks, including molecular graphic convolution neural network (MG-CNN), molecular graph transformer neural network, and atom interaction neural networks. The results of the performance comparison of different feature extractors demonstrate that the MG-CNN model produces the highest accuracy and thus may be most suitable for the SOR prediction. It is hoped that the proposed method can be applied in practice as an auxiliary tool of GCO for the sensory evaluation of key compounds in food ingredients. HighlightsO_LIDifferent deep neural networks were used to predict categorized odor descriptors. C_LIO_LIEnd-to-end-based representation learning was performed for molecular feature extraction. C_LIO_LIMolecular graphs with pre-training of convolution neural networks was most accurate. C_LI Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=82 SRC="FIGDIR/small/488977v2_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@10495a2org.highwire.dtl.DTLVardef@1fbf27eorg.highwire.dtl.DTLVardef@1ed3664org.highwire.dtl.DTLVardef@8e1383_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.