Protein family annotation for the Unified Human Gastrointestinal Proteome by DPCfam clustering
Barone, F.; Russo, E. T.; Villegas Garcia, E. N.; Punta, M.; Cozzini, S.; Ansuini, A.; Cazzaniga, A.
Show abstract
Technological advances in massively parallel sequencing have led to an exponential growth in the number of known protein sequences. Much of this growth originates from metagenomic projects producing new sequences from environmental and clinical samples. The Unified Human Gastrointestinal Proteome (UHGP) catalogue is one of the most relevant metagenomic datasets with applications ranging from medicine to biology. However, the lack of sequence annotation impairs its usability. This work aims to produce a family classification of UHGP sequences to facilitate downstream structural and functional annotation. This is achieved through the release of the DPCfam-UHGP50 dataset containing 10,778 putative protein families generated using DPCfam clustering, an unsupervised pipeline grouping sequences into multi-domain architectures. DPCfam-UHGP50 considerably improves family coverage at protein and residue levels compared to the manually curated repository Pfam. It is our hope that DPCfam-UHGP50 will foster future discoveries in the field of metagenomics of the human gut by the release of a FAIR-compliant database easily accessible via a searchable web server and Zenodo repository.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models 96%
- Clustering FunFams using sequence embeddings improves EC purity 94%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 94%
Similar papers in this journal
- ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM 95%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 94%
- PPanGGOLiN: Depicting microbial diversity via a Partitioned Pangenome Graph 94%
Similar papers in this journal
- AutoPhy: Automated phylogenetic identification of novel protein subfamilies 95%
- Phylogenetic profiling in eukaryotes: The effect of species, orthologous group, and interactome selection on protein interaction prediction 94%
- ProteinWeaver: A Webtool to Visualize Ontology-Annotated Protein Networks 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.