Back

CD-GPT As a Biological Foundation Model Bridging the Gap between Molecular Sequences Through Central Dogma

Zhu, X.; Qin, C.; Wang, F.; Yang, F.; He, B.; Zhao, Y.; Yao, J.

2025-02-11 bioinformatics
10.1101/2024.06.24.600337 bioRxiv
Show abstract

The central dogma serves as a fundamental framework for understanding the flow and expression of genetic information within living organisms, facilitating the connection of diverse biological sequences across molecule types. In this study, we present CD-GPT (Central Dogma Generative Pretrained Transformer), a generative biological foundation model with 1 billion parameters, aiming to capture the sequence relationships between DNA, RNA, and proteins. We model sequences in a unified representational space and employ a shared, multi-molecule vocabulary to narrow their distances in the embedding space effectively. Through extensive pretraining on nucleotide and amino acid sequence data, CD-GPT exhibits exceptional performance in a wide range of predictive and generative downstream tasks, including mono-molecular and multi-molecular analyses. Notably, CD-GPT excels in tasks such as genomic element detection, protein property prediction, RNA-protein interaction identification and also generative tasks like protein generation and reverse translation. The versatility of CD-GPT opens up promising avenues for advanced multi-omics analysis.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.