Back

enviLink: A database linking contaminant biotransformation rules to enzyme classes in support of functional association mining

Schmid, E.; Fenner, K.

2021-05-21 bioinformatics
10.1101/2021.05.20.442588 bioRxiv
Show abstract

MotivationThe ability to assess and engineer biotransformation of chemical contaminants present in the environment requires knowledge on which enzymes can catalyze specific contaminant biotransformation reactions. For the majority of over 100000 chemicals in commerce such knowledge is not available. Enumeration of enzyme classes potentially catalyzing observed or de novo predicted contaminant biotransformation reactions can support research that aims at experimentally uncovering enzymes involved in contaminant biotransformation in complex natural microbial communities. DatabaseenviLink is a new data module integrated into the enviPath database and contains 316 theoretically derived linkages between generalized biotransformation rules used for contaminant biotransformation prediction in enviPath and 3rd level EC classes. Rule-EC linkages have been derived using two reaction databases, i.e., Eawag-BBD in enviPath, focused on contaminant biotransformation reactions, and KEGG. 32.6% of identified rule-EC linkages overlap between the two databases, whereas 40.2% and 27.2%, respectively, are originating from Eawag-BBD and KEGG only. Implementation and availabilityenviLink is encoded in RDF triples as part of the enviPath RDF database. enviPath is hosted on a public webserver (envipath.org) and all data is freely available for non-commercial use. enviLink can be searched online for individual transformation rules of interest (https://tinyurl.com/y63ath3k) and is also fully downloadable from the supporting materials (i.e., Jupyter notebook "enviLink" and tsv files provided through GitHub at https://github.com/emanuel-schmid/enviLink).

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.