NetSyn: genomic context exploration of protein families
Stam, M.; Langlois, j.; Chevalier, C.; Reboul, G.; Bastard, K.; Medigue, C.; Vallenet, D.
Show abstract
BackgroundThe growing availability of large genomic datasets presents an opportunity to discover new metabolic pathways and enzymatic reactions useful for industrial or synthetic biological applications. Efforts to identify new enzyme functions in this huge number of sequences cannot be achieved without the help of bioinformatics tools and the development of new strategies. Standard methods for assigning a biological function to a gene are based on sequence similarity. However, another way is to mine databases to identify conserved gene clusters (i.e. syntenies). Indeed in prokaryotic genomes, genes involved in the same pathway are frequently encoded in a single locus with an operonic organisation. This Genomic Context (GC) conservation is considered as a reliable indicator of functional relationships, and thus is a promising approach to improve the gene function prediction. MethodsHere we present NetSyn (Network Synteny), a tool which aims to cluster protein sequences according to the similarity of their genomic context rather than their sequence similarity. Starting from a set of protein sequences of interest, NetSyn retrieves neighbouring genes from the corresponding genomes as well as their protein sequences. Homologous protein families are then computed to measure GC conservation between each pair of input sequences using a GC conservation score. A network is created where nodes represent the input proteins and edges the fact that two proteins share a common GC. The weight of the edges corresponds to the synteny conservation score. The network is then partitioned into clusters of proteins sharing a high degree of synteny conservation. ResultsAs a proof of concept, we used NetSyn on two different datasets. The first one is made of homologous sequences of an enzyme family (the BKACE family, previously named DUF849) to divide it into sub-families of specific activities. NetSyn was able to go a step further by providing additional subfamilies to those already published. The second dataset corresponds to a set of non-homologous proteins consisting of different Glycosyl Hydrolases (GH) with the aim of interconnecting them and finding conserved operon-like genomic structures. NetSyn was able to find a locus composed of 3 non homologous glycosyl hydrolases, involved in the degradation of xyloglucan, in 162 genomes. DiscussionNetSyn is able to cluster proteins according to their genomic context, making it possible to establish functional links between proteins without considering their sequence similarity alone. We showed that NetSyn is efficient in exploring large protein families to define iso-functional groups. It can also highlight functional interactions between proteins from different families and predicts new conserved genomic structures that have not yet been experimentally characterised. NetSyn can also be useful for pinpointing annotation errors that have been propagated across databases, and for suggesting annotations on proteins currently annotated as "unknown". NetSyn is freely available at https://github.com/labgem/netsyn.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- panRGP: a pangenome-based method to predict genomic islands and explore their diversity 97%
- DeNoFo: a file format and toolkit for standardised, comparable de novo gene annotation 95%
- No one tool to rule them all: Prokaryotic gene prediction tool performance is highly dependent on the organism of study 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.