Back

Supervised and Unsupervised Classification of Cocoa Bean Origin and Processing using Liquid Chromatography-Mass Spectrometry

Kumar, S.; D'Souza, R. N.; Behrends, B.; Corno, M.; Ullrich, M. S.; Kuhnert, N.; Huett, M.-T.

2020-02-10 bioinformatics
10.1101/2020.02.09.940577 bioRxiv
Show abstract

Liquid Chromatography-Mass Spectrometry (LC-MS) provides an unprecedented wealth of metabolomics information for food products, including insights into compositional changes during food processing. Here, we employed the largest available LC-MS dataset of around 300 cocoa bean samples to assess the capability of two popular multivariate classification methods, principal component analysis (PCA) and linear decomposition analysis (LDA), for studying bean geographic origin and responsible characteristic compounds. The unsupervised method, PCA, only provides a limited separation in bean origin. Expectedly, the supervised method, LDA, provides a better origin clustering. However, it suffers from a strong, nonlinear dependence on the set of compounds used in the analysis. We show that for LDA a compound filtering criterion based on Gaussian intensity distributions dramatically enhances origin clustering of samples, thus increasing its predictive efficiency. In this form, the supervised method of LDA holds the possibility to identify potential markers of a specific origin.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.