Back

PROTEORIZER: A holistic approach to untangle functional consequences of variants of unknown significance.

Schmenger, T.; Diwan, G.; Russell, R. B.

2024-07-19 bioinformatics
10.1101/2024.07.16.603688 bioRxiv
Show abstract

Most in silico tools only use data closely related to the gene-of-interest or initial research question. This gene-focused research is prone to ignoring low-count and rare variants in the same or similar genes, even if available informational could be sufficient to deduce functional consequences by combining knowledge from many similar genes. Proteorizer is a web tool that aims to bridge the gap between protein-centric knowledge and the functional context this knowledge creates. We use curated and reviewed data from UniProt to collect available residue information for the queried protein as well as orthologs. By defining functional clusters based on intramolecular distances of residues with available functional information it is possible to use these to extrapolate the effect of a VUS solely based on known functions of nearby residues, hence contextualizing the variant with pre-existing knowledge. We show that pathogenic variants are more likely to be a part of functional hotspots and present several case studies (ALPP p.Ser244Gly, CANT1 p.Ile171Phe, ARL3 p.Tyr90Cys, IL6R p.His280Pro and RAF1 p.Ser259Ala) to highlight the applicability and usefulness of this approach. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=121 SRC="FIGDIR/small/603688v1_ufig1.gif" ALT="Figure 1"> View larger version (26K): org.highwire.dtl.DTLVardef@a1d75forg.highwire.dtl.DTLVardef@142c339org.highwire.dtl.DTLVardef@1f1161org.highwire.dtl.DTLVardef@1adfe85_HPS_FORMAT_FIGEXP M_FIG C_FIG Proteorizer is an explorative tool that takes variants from laboratory or clinical settings and contextualizes the variants based on prior information from the protein of interest and similar proteins according to where these functional positions are located in the 3D structure of the protein of interest.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.