GENA-Web - GENomic Annotations Web Inference using DNA language models
Shmelev, A.; Petrov, M.; Penzar, D.; Akhmetyanov, N.; Tavritskiy, M.; Mamontov, S.; Kuratov, Y.; Burtsev, M.; Kardymon, O.; Fishman, V.
Show abstract
The advent of advanced sequencing technologies has significantly reduced the cost and increased the feasibility of assembling high-quality genomes. Yet, the annotation of genomic elements remains a complex challenge. Even for species with comprehensively annotated reference genomes, the functional assessment of individual genetic variants is not straightforward. In response to these challenges, recent breakthroughs in machine learning have led to the development of DNA language models. These transformer-based architectures are designed to tackle a wide array of genomic tasks with enhanced efficiency and accuracy. In this context, we introduce GENA-Web, a web-based platform that consolidates a suite of genome annotation tools powered by DNA language models. The version of GENA-Web presented here encompasses a diverse set of models trained on human data, including the prediction of promoter activity, annotation of splice sites, determination of various chromatin features, and a model for scoring of enhancer activity in Drosophila. GENA-Web is accessible online at https://dnalm.airi.net/
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins 95%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 95%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 94%
Similar papers in this journal
- Sensitive and error-tolerant annotation of protein-coding DNA with BATH 95%
- CanDrivR-CS: A Cancer-Specific Machine Learning Framework for Distinguishing Recurrent and Rare Variants 94%
- WAS IT A MATch I SAW? Approximate palindromes lead to overstated false match rates in benchmarks using reversed sequences 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.