DomainRank: Improving Biological Data Sets With Domain Knowledge and Google's PageRank
Schneider, M.; Rappsilber, J.; Brock, O.
Show abstract
MotivationThe quality of biological data crucially affects progress in science. This quality can be improved with better measurement devices, more sophisticated experimental designs, or repetitious measurements. Each of these options is associated with substantial costs. We present a simple computational tool as an alternative. This algorithmic tool, called DomainRank, leverages simple domain knowledge and overlapping information content in biological network data to improve measurement quality at a negligible cost. Following the simple computational template of Domain-Rank, researchers can boost the confidence of their own data with little effort. ResultsWe demonstrate the performance of DomainRank in three test cases: DomainRank finds 14.9% more interactions in quantitative proteomics experiments, improves the precision of predicted residue-residue contacts from co-evolutionary data by up to 11.6% (averaged over 882 proteins), and identifies 89.2% more cross-links in photo-crosslinking/mass spectrometry (photo-CLMS) experiments. Although our proposed template is specialized on biological network data, we view this approach as an universal computational tool for data improvement that could be routinely applied in many disciplines. AvailabilityAn implementation of DomainRank is freely available: https://github.com/Rappsilber-Laboratory/pagerank-refine
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 96%
- Limits and potential of combined folding and docking using PconsDock. 95%
- Patch-DCA: Improved Protein Interface Prediction by utilizing Structural Information and Clustering DCA scores 95%
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 96%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 95%
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 94%
Similar papers in this journal
Similar papers in this journal
- PrePPI: A structure informed proteome-wide database of protein-protein interactions 96%
- Prediction of disordered regions in proteins with recurrent Neural Networks and protein dynamics 96%
- Integrating multimeric threading with high-throughput experiments for structural interactome of Escherichia coli 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.