Streamlined use of protein structures in variant analysis
Kaur, S.; Sikta, N.; Schafferhans, A.; Bordin, N.; Cowley, M. J.; Thomas, D. M.; Ballinger, M. L.; O'Donoghue, S. I.
Show abstract
MotivationVariant analysis is a core task in bioinformatics that requires integrating data from many sources. This process can be helped by using 3D structures of proteins, which can provide a spatial context that can provide insight into how variants affect function. Many available tools can help with mapping variants onto structures; but each has specific restrictions, with the result that many researchers fail to benefit from valuable insights that could be gained from structural data. ResultsTo address this, we have created a streamlined system for incorporating 3D structures into variant analysis. Variants can be easily specified via URLs that are easily readable and writable, and use the notation recommended by the Human Genome Variation Society (HGVS). For example, https://aquaria.app/SARS-CoV-2/S/?N501Y specifies the N501Y variant of SARS-CoV-2 S protein. In addition to mapping variants onto structures, our system provides summary information from multiple external resources, including COSMIC, CATH-FunVar, and PredictProtein. Furthermore, our system identifies and summarizes structures containing the variant, as well as the variant-position. Our system supports essentially any mutation for any well-studied protein, and uses all available structural data -- including models inferred via very remote homology -- integrated into a system that is fast and simple to use. By giving researchers easy, streamlined access to a wealth of structural information during variant analysis, our system will help in revealing novel insights into the molecular mechanisms underlying protein function in health and disease. AvailabilityOur resource is freely available at the project home page (https://aquaria.app). After peer review, the code will be openly available via a GPL version 2 license at https://github.com/ODonoghueLab/Aquaria. PSSH2, the database of sequence-to-structure alignments, is also freely available for download at https://zenodo.org/record/4279164. Contactsean@odonoghuelab.org Supplementary informationNone.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Fast and scalable querying of eukaryotic linear motifs with gget elm 96%
- AlphaPulldown - a Python package for protein-protein interaction screens using AlphaFold-Multimer 95%
- MutaFrame - an interpretative visualization framework for deleteriousness prediction of missense variants in the human exome 94%
Similar papers in this journal
- dms-viz: Structure-informed visualizations for deep mutational scanning and other mutation-based datasets 94%
- alignparse: A Python package for parsing complex features from high-throughput long-read sequencing 92%
- diverse-seq: an application for alignment-free selecting and clustering biological sequences 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.