Language model generates cis-regulatory elements across prokaryotes
Xia, Y.; Sun, J.; Du, X.; Liang, Z.; Shi, W.; Guo, S.; Huo, Y.-X.
Show abstract
Deep learning had succeeded in designing Cis-regulatory elements (CREs) for certain species, but necessitated training data derived from experiments. Here, we present Promoter-Factory, a protocol that leverages language models (LM) to design CREs for prokaryotes without experimental prior. Millions of sequences were drawn from thousands of prokaryotic genomes to train a suite of language models, named PromoGen2, and achieved the highest zero-shot promoter strength prediction accuracy among tested LMs. Artificial CREs designed with Promoter-Factory achieved a 100% success rate to express gene in Escherichia coli, Bacillus subtilis, and Bacillus licheniformis. Furthermore, most of the promoters designed targeting Jejubacter sp. L23, a halophilic bacterium without available CREs, were active and successfully drove lycopene overproduction. The generation of 2 million putative promoters across 1,757 prokaryotic genera, along with the Promoter-Factory protocol, will significantly expand the sequence space and facilitate the development of an extensive repertoire of prokaryotic CREs.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NanoLoop: A deep learning framework leveraging Nanopore sequencing for chromatin loop prediction 94%
- Automatically Defining Protein Words for Diverse Functional Predictions Based on Attention Analysis of a Protein Language Model 94%
- DeDoc2 identifies and characterizes the hierarchy and dynamics of chromatin TAD-like domains in the single cells 94%
Similar papers in this journal
- High-precision cell-type mapping and annotation of single-cell spatial transcriptomics with STAMapper 94%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 94%
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.