What is hidden in the darkness? Deep-learning assisted large-scale protein family curation uncovers novel protein families and folds
Durairaj, J.; Waterhouse, A.; Mets, T.; Brodiazhenko, T.; Abdullah, M.; Akdel, M.; Andreeva, A.; Bateman, A.; Hauryliuk, V.; Tenson, T.; Schwede, T.; Pereira, J.
Show abstract
Driven by the development and upscaling of fast genome sequencing and assembly pipelines, the number of protein-coding sequences deposited in public protein sequence databases is increasing exponentially. Recently, the dramatic success of deep learning-based approaches applied to protein structure prediction has done the same for protein structures. We are now entering a new era in protein sequence and structure annotation, with hundreds of millions of predicted protein structures made available through the AlphaFold database. These models cover most of the catalogued natural proteins, including those difficult to annotate for function or putative biological role based on standard, homology-based approaches. In this work, we quantified how much of such "dark matter" of the natural protein universe was structurally illuminated by AlphaFold2 and modelled this diversity as an interactive sequence similarity network that can be navigated at https://uniprot3d.org/atlas/AFDB90v4. In the process, we discovered multiple novel protein families by searching for novelties from sequence, structure, and semantic perspectives. We added a number of them to Pfam, and experimentally demonstrate that one of these belongs to a novel superfamily of translation-targeting toxin-antitoxin systems, TumE-TumA. This work highlights the role of large-scale, evolution-driven protein comparison efforts in combination with structural similarities, genomic context conservation, and deep-learning based function prediction tools for the identification of novel protein families, aiding not only annotation and classification efforts but also the curation and prioritisation of target proteins for experimental characterisation.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Purging genomes of contamination eliminates systematic bias from evolutionary analyses of ancestral genomes 94%
- Surface frustration re-patterning underlies the structural landscape and evolvability of fungal orphan candidate effectors 94%
- High-resolution cryo-EM structure of urease from the pathogen Yersinia enterocolitica 94%
Similar papers in this journal
- Bridging themes: short protein segments found in different architectures 95%
- Newly developed structure-based methods do not outperform standard sequence-based methods for large-scale phylogenomics 95%
- Exploring large protein sequence space through homology- and representation-based hierarchical clustering 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.