CD-GPT As a Biological Foundation Model Bridging the Gap between Molecular Sequences Through Central Dogma
Zhu, X.; Qin, C.; Wang, F.; Yang, F.; He, B.; Zhao, Y.; Yao, J.
Show abstract
The central dogma serves as a fundamental framework for understanding the flow and expression of genetic information within living organisms, facilitating the connection of diverse biological sequences across molecule types. In this study, we present CD-GPT (Central Dogma Generative Pretrained Transformer), a generative biological foundation model with 1 billion parameters, aiming to capture the sequence relationships between DNA, RNA, and proteins. We model sequences in a unified representational space and employ a shared, multi-molecule vocabulary to narrow their distances in the embedding space effectively. Through extensive pretraining on nucleotide and amino acid sequence data, CD-GPT exhibits exceptional performance in a wide range of predictive and generative downstream tasks, including mono-molecular and multi-molecular analyses. Notably, CD-GPT excels in tasks such as genomic element detection, protein property prediction, RNA-protein interaction identification and also generative tasks like protein generation and reverse translation. The versatility of CD-GPT opens up promising avenues for advanced multi-omics analysis.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 96%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 96%
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 95%
Similar papers in this journal
- Sampling from Disentangled Representations of Single-Cell Data Using Generative Adversarial Networks 96%
- Learning latent embedding of multi-modal single cell data and cross-modality relationship simultaneously 95%
- PlasRAG: comprehensive plasmid characterization and retrieval through sequence-text alignment 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.