Unified Genomic and Chemical Representations Enable Bidirectional Bio-synthetic Gene Cluster and Natural Product Retrieval
Liu, G.; Li, Y.; Ong, G.; Wong, F. T.; Tay, D. W. P.; Lim, Y. H.; Foo, C. S.; Koh, L. C. W.
Show abstract
Natural product discovery is increasingly driven by the ability to analyze microbial genomes for biosynthetic gene clusters (BGCs) that encode secondary metabolites. While existing approaches have successfully linked BGCs to broad classes of chemical products, they typically operate in a single modality (genomic or chemical) limiting the scope of bidirectional prediction. In this work, we propose a multimodal framework that integrates genomic and chemical information by projecting embeddings derived from pretrained language models into a common representation space. We embed genomic sequences using a BGC foundation model and represent molecules through a chemical language model, then use a metric learning model to co-embed BGCs and their associated chemical structures. This co-embedding space allows us to quantify the similarity between BGCs and compounds using similarity measures, enabling both forward and inverse retrieval tasks. Beyond retrieval, we show that the shared space can guide strain selection for targeted compounds. By identifying BGCs closest to a query compound in the embedding space, we prioritize microbial strains that encode similar clusters, thereby streamlining genome mining and retrobiosynthetic design efforts. This approach represents a generalizable, scalable strategy to bridge biological and chemical modalities in natural product discovery.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cross-Modality and Self-Supervised Protein Embedding for Compound-Protein Affinity and Contact Prediction 96%
- One-Hot News: Drug Synergy Models Shortcut Molecular Features 96%
- DTI-Voodoo: machine learning over interaction networks and ontology-based background knowledge predicts drug-target interactions 95%
Similar papers in this journal
- GexMolGen: Cross-modal Generation of Hit-like Molecules via Large Language Model Encoding of Gene Expression Signatures 96%
- BatchDTA: Implicit batch alignment enhances deep learning-based drug-target affinity estimation 96%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 96%
Similar papers in this journal
- A deep learning framework for high-throughput mechanism-driven phenotype compound screening 95%
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 94%
- TrustAffinity: accurate, reliable and scalable out-of-distribution protein-ligand binding affinity prediction using trustworthy deep learning 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.