Predicting evolutionary rate as a pretraining task improves genome language model representations
Consens, M. E.; Yang, K. K.; Hall, J.; Conard, A. M.; Wang, B.; Crawford, L.; Moses, A.; Lu, A. X.
Show abstract
Genome language models (gLM) have the potential to further understanding of regulatory genomics without requiring labeled data. Most gLMs are pretrained using sequence reconstruction tasks inspired by natural language processing, but recent studies have shown that these gLMs often fail to capture biological signal. To overcome this, we introduce pretraining tasks that predict the rate of evolution. These tasks are designed so that they can be composed with sequence reconstruction, enabling a controlled comparison of predicting sequence only, evolutionary rate only, or both. To address gaps in existing evaluations, we developed a suite of biologically grounded benchmarks. Across these tasks, and for established variant effect prediction benchmarks, models pretrained on both sequence and evolutionary rate outperform those trained on sequence alone, and training on evolutionary rate can make the even the relatively small models in our work competitive with much larger existing gLMs for some tasks. These results establish evolution as a key training target for genome-scale models.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 95%
- CellPhy: accurate and fast probabilistic inference of single-cell phylogenies from scDNA-seq data 95%
- DelSIEVE: cell phylogeny model of single nucleotide variants and deletions from single-cell DNA sequencing data 94%
Similar papers in this journal
- Representation Learning of Genomic Sequence Motifs with Convolutional Neural Networks 96%
- Deep Mendelian Randomization: Investigating the causal knowledge of genomic deep learning models 95%
- Improving deep models of protein-coding potential with a Fourier-transform architecture and machine translation task 95%
Similar papers in this journal
- Vector-clustering Multiple Sequence Alignment: Aligning into the twilight zone of protein sequence similarity with protein language models 96%
- Memory-bound k-mer selection for large evolutionary diverse reference libraries 95%
- Domain adaptive neural networks improvecross-species prediction of transcription factor binding 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.