SegmentNT: annotating the genome at single-nucleotide resolution with DNA foundation models
de Almeida, B. P.; Dalla-Torre, H.; Richard, G.; Blum, C.; Hexemer, L.; Gelard, M.; Pandey, P.; Laurent, S.; Laterre, A.; Lang, M.; Sahin, U.; Beguir, K.; Pierrot, T.
Show abstract
Genome annotation models that directly analyze DNA sequences are indispensable for modern biological research, enabling rapid and accurate identification of genes and other functional elements. This capability is paramount as the volume of sequenced genomes rapidly expands, making the need for efficient and accurate annotation methods increasingly critical, particularly in the context of genetic variant prediction and in-silico sequence design. Current annotation tools are typically developed for specific element classes and trained from scratch using supervised learning on datasets that are often limited in size. This approach constrains their performance and ability to generalize to new genomes. Here, we frame the genome annotation problem as instance segmentation and introduce a novel methodology for fine-tuning pre-trained DNA foundation models to segment 14 different genic and regulatory elements at single-nucleotide resolution. We leverage the self-supervised pre-trained model Nucleotide Transformer (NT) to develop a general segmentation model, SegmentNT, capable of processing DNA sequences up to 50kb long. By utilizing pre-trained weights from NT, SegmentNT surpasses the performance of several ablation models and baselines, including convolutional networks with one-hot encoded nucleotide sequences and large models trained from scratch. We demonstrate state-of-the-art performance on gene annotation, splice site and regulatory elements detection throughout the genome. We also leveraged our framework to accommodate two extra DNA foundation models, Enformer and Borzoi, extending the sequence context up to 500kb and enhancing performance on regulatory elements. Finally, we show that a SegmentNT model trained on human genomic elements generalizes to elements of different species, and a multi-species SegmentNT model achieves strong generalization across unseen species. Our approach is readily extensible to additional genomic elements and species. We have made our SegmentNT human and multi-species models, as well as the SegmentEnformer and SegmentBorzoi models, available on our github repository in Jax and HuggingFace space in Pytorch.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 96%
- Epiphany: predicting Hi-C contact maps from 1D epigenomic signals 96%
- Towards In-Silico CLIP-seq: Predicting Protein-RNA Interaction via Sequence-to-Signal Learning 96%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 96%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 96%
- Epitome: Predicting epigenetic events in novel cell types with multi-cell deep ensemble learning 95%
Similar papers in this journal
- The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics 98%
- Haplotype-aware variant calling enables high accuracy in nanopore long-reads using deep neural networks 96%
- scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.