Back

Varia: Prediction, analysis and visualisation of variable genes

Mackenzie, G.; Jensen, R.; Lavstsen, T.; Otto, T.

2020-12-16 bioinformatics
10.1101/2020.12.15.422815 bioRxiv
Show abstract

Assessing the diversity or expression of variable gene families in pathogens can inform about immune escape mechanisms or host interaction phenotypes of clinical relevance. However, obtaining the sequences and quantifying their expression is a challenge. Here, we present a tool, which based on unique sequence tag similarity between members of a gene family, predicts the domains encoded by the queried gene. As an example, we are using the var gene family, encoding the major virulence proteins (PfEMP1) of the human malaria parasite, Plasmodium falciparum. We developed Varia, which predicts the likely var gene sequence and encoded protein domain composition of a gene from short sequence tags. We provide a new extended annotated var genome database, in which Varia identifies genes with identical tag sequences and compares these to return the most probable domain composition of the query gene. Varias ability to predict correct PfEMP1 domain compositions from short var sequence tags was tested in two complementary pipelines to (a) return the putative gene sequences and domain compositions of the query gene from any partial sequence provided, thereby enabling detailed assessment of specific genes putative function and experimental validation of these (b) to accommodate rapid profiling of var gene expression in complex patient samples, by compiling the overall domain prevalence among var transcripts predicted identified and quantified by next generation sequencing of so-called var DBL-sequence tags. Availability and implementationVaria is available on GitHub (https://github.com/GCJMacken-zie/Varia) under the MIT license. Contactthomasl@sund.ku.dk, thomasdan.otto@glasgow.ac.uk

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

1
BMC Bioinformatics
457 papers in training set
Top 0.1%
39.6%
2
BMC Genomics
406 papers in training set
Top 0.3%
10.6%
50% of probability mass above
3
Bioinformatics
1204 papers in training set
Top 3%
9.8%
4
Bioinformatics Advances
203 papers in training set
Top 0.9%
5.5%
5
Nucleic Acids Research
1281 papers in training set
Top 5%
3.4%
6
Scientific Reports
3612 papers in training set
Top 33%
3.2%
7
GigaScience
212 papers in training set
Top 1%
2.8%
8
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.7%
9
PLOS ONE
5266 papers in training set
Top 51%
1.5%
10
Malaria Journal
58 papers in training set
Top 0.6%
1.5%
11
PLOS Computational Biology
1863 papers in training set
Top 16%
1.3%
12
PLOS Genetics
862 papers in training set
Top 8%
1.3%
13
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
14
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
15
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.1%
16
F1000Research
88 papers in training set
Top 3%
1.1%
17
Peer Community Journal
281 papers in training set
Top 4%
1.0%
18
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.8%
19
Communications Biology
993 papers in training set
Top 30%
0.8%
20
Genome Medicine
183 papers in training set
Top 6%
0.6%
21
Genome Biology
637 papers in training set
Top 9%
0.6%
22
Wellcome Open Research
67 papers in training set
Top 2%
0.6%
23
Database
61 papers in training set
Top 1%
0.6%