Back

tbg - a new file format for genomic data

Gerber, S.; Pfenninger, M.; Schell, T.; Schoennenbeck, P.

2021-03-16 genomics
10.1101/2021.03.15.435393 bioRxiv
Show abstract

MotivationThe question of determining whether a Single-Nucleotide Polymorphism (SNP) or a variant in general leads to a change in the amino acid sequence of a protein coding gene is often a laborious and time-consuming challenge. Here, we introduce the tbg file format for storing genomic data and tbg-tools, a user-friendly toolbox for the faster analysis of SNPs. The file format stores information for each nucleotide in each gene, allowing to predict which change in the amino acid sequence will be caused by a variant in the nucleotide sequence. Our new tool therefore has the potential to make biological sense of the unprecedented amount of genome-wide genetic variation that researchers currently face. ResultsThe new tab-separated file for storing the nucleotide data can be easily analyzed and used for a wide variety of biological research. It is also possible to automate some of these analyses using the additional analysis tools from tbg-tools Availabilitytbg-tools is written in Python and allows the installation from the command line. It can be found on https://github.com/Croxa/tbg-tools. Contactpschoenn@students.uni-mainz.de

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.