Back

AlphaFind: Discover structure similarity across the entire known proteome

Prochazka, D.; Slaninakova, T.; Olha, J.; Rosinec, A.; Gresova, K.; Janosova, M.; Cillik, J.; Porubska, J.; Svobodova, R.; Dohnal, V.; Antol, M.

2024-02-18 bioinformatics Community evaluation
10.1101/2024.02.15.580465 bioRxiv
Show abstract

AlphaFind is a web-based search engine that provides fast structure-based retrieval in the entire set of AlphaFold DB structures. Unlike other protein processing tools, AlphaFind is focused entirely on tertiary structure, automatically extracting the main 3D features of each protein chain and using a machine learning model to find the most similar structures. This indexing approach and the 3D feature extraction method used by AlphaFind have both demonstrated remarkable scalability to large datasets as well as to large protein structures. The web application itself has been designed with a focus on clarity and ease of use. The searcher accepts any valid Uniprot ID, PDB ID or gene symbol as input, and returns a set of similar protein chains from AlphaFold DB, including various similarity metrics between the query and each of the retrieved results. In addition to the main search functionality, the application provides 3D visualizations of protein structure superpositions in order to allow researchers to instantly analyze the structural similarity of the retrieved results. The AlphaFind web application is available online for free and without any registration at https://alphafind.fi.muni.cz. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=82 SRC="FIGDIR/small/580465v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@1bf5ae8org.highwire.dtl.DTLVardef@1e975d7org.highwire.dtl.DTLVardef@378598org.highwire.dtl.DTLVardef@123e439_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.