Back

GENA-Web - GENomic Annotations Web Inference using DNA language models

Shmelev, A.; Petrov, M.; Penzar, D.; Akhmetyanov, N.; Tavritskiy, M.; Mamontov, S.; Kuratov, Y.; Burtsev, M.; Kardymon, O.; Fishman, V.

2024-04-29 bioinformatics
10.1101/2024.04.26.591391 bioRxiv
Show abstract

The advent of advanced sequencing technologies has significantly reduced the cost and increased the feasibility of assembling high-quality genomes. Yet, the annotation of genomic elements remains a complex challenge. Even for species with comprehensively annotated reference genomes, the functional assessment of individual genetic variants is not straightforward. In response to these challenges, recent breakthroughs in machine learning have led to the development of DNA language models. These transformer-based architectures are designed to tackle a wide array of genomic tasks with enhanced efficiency and accuracy. In this context, we introduce GENA-Web, a web-based platform that consolidates a suite of genome annotation tools powered by DNA language models. The version of GENA-Web presented here encompasses a diverse set of models trained on human data, including the prediction of promoter activity, annotation of splice sites, determination of various chromatin features, and a model for scoring of enhancer activity in Drosophila. GENA-Web is accessible online at https://dnalm.airi.net/

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.