L2G: Repurposing Language Models for Genomics Tasks
Cheng, W.; Shen, J.; Khodak, M.; Ma, J.; Talwalkar, A.
Show abstract
Pre-trained language models have transformed the field of natural language processing (NLP), and their success has inspired efforts in genomics to develop domain-specific foundation models (FMs). However, creating high-quality genomic FMs from scratch is resource-intensive, requiring significant computational power and high-quality pre-training data. The success of large language models (LLMs) in NLP has largely been driven by industrial-scale efforts leveraging vast, diverse corpora and massive computing infrastructure. In this work, we aim to bypass the data and computational bottlenecks of creating genomic FMs from scratch and instead propose repurposing existing LLMs for genomics tasks. Inspired by the recently observed cross-modal transfer phenomenon - where transformers pre-trained on natural language can generalize to other modalities - we introduce L2G, which adapts a pre-trained LLM architecture for genomics using neural architecture search (NAS) and a novel three-stage training procedure. Remarkably, without requiring extensive pre-training on DNA sequence data, L2G achieves superior performance to fine-tuned genomic FMs and task-specific models on more than half of tasks across multiple genomics benchmarks. In an enhancer activity prediction task, L2G further demonstrates its capacity to identify significant transcription factor motifs. Our work not only highlights the generalizability and efficacy of language models in out-of-domain tasks such as genomics, but also opens new avenues for more efficient and less resource-intensive methodologies in genomic research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Graph Contrastive Learning of Subcellular-resolution Spatial Transcriptomics Improves Cell Type Annotation and Reveals Critical Molecular Pathways 95%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 95%
- Deep generative modeling and clustering of single cell Hi-C data 95%
Similar papers in this journal
- NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site prediction 97%
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 96%
- EvoAug-TF: Extending evolution-inspired data augmentations for genomic deep learning to TensorFlow 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.