Predicting Molecular Taste: Multi-Label and Multi-Class Classification
Ramanathan, V.; DN, S. S.
Show abstract
Predicting the taste of chemical compounds is a complex task and has been a challenge for decades. This study explores the application of machine learning to predict taste profiles of chemical compounds using the ChemTastesDB dataset, comprising 2,944 tastants categorized into 44 taste labels and 9 taste classes. Addressing the challenges of label imbalance and correlation, the dataset was preprocessed using iterative stratified sampling and feature representations such as Mordred descriptors, Morgan fingerprints, and Daylight fingerprints. Baseline random forest models, along with binary relevance and classifier chains, were employed for multi-label classification, with evaluation metrics including micro-averaged F1 scores, precision, and recall. Results demonstrated that binary relevance models, particularly with Morgen fingerprints, achieved superior F1 scores, outperforming classifier chains likely due to random label ordering. Label correlation analysis via co-occurrence matrices and community detection revealed significant associations between taste labels, providing deeper insights into molecular taste interactions. Feature importance analysis highlighted structural elements influencing taste prediction. This work underscores the potential of computational models in advancing flavor science and paves the way for future exploration with deep learning and optimized label dependencies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DenovoProfiling: a webserver for de novo generated molecule library profiling 93%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 93%
- DrugForm-DTA: Towards real-world drug-target binding Affinity Model 92%
Similar papers in this journal
- Pred-AHCP: Robust feature selection enabled Sequence Specific Prediction of Anti-Hepatitis C Peptides via Machine Learning 94%
- Predicting Antimicrobial Activity for Untested Peptide-Based Drugs Using Collaborative Filtering and Link Prediction 94%
- Identification of Family-Specific Features in Cas9 and Cas12 Proteins: A Machine Learning Approach Using Complete Protein Feature Spectrum 93%
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
- PharmaNet: Pharmaceutical discovery with deep recurrent neural networks. 93%
- Identification of Natural Antiviral Drug Candidates Against Tilapia Lake Virus: Computational Drug Design Approaches 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.