OmniNA: A foundation model for nucleotide sequences
Shen, X.; Li, X.
10.1101/2024.01.14.575543 bioRxivShow abstract
Foundation models have demonstrated exceptional efficacy across diverse downstream tasks. However, within the realms of genomics and transcriptomics, a notable gap persists in the availability of models that afford a comprehensive understanding of nucleotide sequence principles across various species. Here, we present OmniNA, a foundation generative model designed for comprehensive nucleotide sequence learning. The model was pre-trained on 91.7 million nucleotide sequences and the corresponding annotations encompassing 1076.2 billion bases and 197 million words spanning a multitude of species. We demonstrated OmniNA gains the capacity to understand the semantics of the nucleotide sequence and textual annotations by analyzing the learned representation of the pre-trained model. OmniNA can be fine-tuned to align multiple nucleotide learning tasks with natural language paradigms. We demonstrate OmniNA-1.7B surpasses or rivals state-of-the art methods in 17 nucleotide tasks, encompassing nucleotide sequences detection and species classification. The models understanding of nucleotide grammars enhances its capability to reveal the mutation effect of nucleotide sequence on DNA and RNA processing. We hereby release the OmniNA-1.7B model as an open-source contribution to the research community. This foundation model signifies a step toward advancing our comprehension of nucleotide sequences across diverse species and holds substantial promise to facilitating genomics and transcriptomics research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 97%
- Correcting gradient-based interpretations of deep neural networks for genomics 96%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 96%
Similar papers in this journal
- EvoAug-TF: Extending evolution-inspired data augmentations for genomic deep learning to TensorFlow 96%
- SMILE: Mutual Information Learning for Integration of Single Cell Omics Data 96%
- DeepLocRNA: An Interpretable Deep Learning Model for Predicting RNA Subcellular Localization with domain-specific transfer-learning 96%
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 98%
- Probabilistic tensor decomposition extracts better latent embeddings from single-cell multiomic data 96%
- Coralysis enables sensitive identification of imbalanced cell types and states in single-cell data via multi-level integration 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.