TB-annotator: a scalable web application that allows in-depth analysis of very large sets of publicly available Mycobacterium tuberculosis complex genomes
Senelle, G.; Guyeux, C.; Refregier, G.; Sola, C.
Show abstract
Tuberculosis continues to be one of the most threatening bacterial diseases in the world. However, we currently have more than 160,000 Short Read Archives (SRAs) of Mycobacterium tuberculosis complex. Such a large amount of data should help to the understanding and the fight against this bacterium. To accomplish this, it would be necessary to thoroughly and comprehensively examine this significant mass of data. This is what TB-Annotator proposes to do, combining a database containing all the diversity of these 160,000 SRAs (at least, SRAs with a reasonable read size and quality), and a fully featured analysis platform to explore and query such a large amount of data. The objective of this article is to present this platform centered on the key notion of exclusivity, to show its numerous capacities (detection of single nucleotide variants, insertion sequences, deletion regions, spoligotyping, etc.) and its general functioning. We will compare TB-Annotator to existing tools for the study of tuberculosis, and show that its objectives are original and have no equivalent at present. The database on which it is based will be presented, with the numerous advanced search queries and screening capacities it offers, and the interest and originality of its phylogenetic tree navigation interface will be detailed. We will end this article with examples of the achievements made possible by the TB-Annotator, followed by avenues for future improvement.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning-based approach KEVOLVE efficiently identifies SARS-CoV-2 variant-specific genomic signatures 95%
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 95%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 94%
Similar papers in this journal
Similar papers in this journal
- BacAnt: A Combination Annotation Server for Bacterial DNA Sequences to Identify Antibiotic Resistance Genes, Integrons, and Transposable Elements. 95%
- DAnIEL: A User-Friendly Web Server for Fungal ITS Amplicon Sequencing Data 94%
- Using Deep Learning for Gene Detection and Classification in Raw Nanopore Signals 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.