Multifaceted quality assessment of gene repertoire annotation with OMArk
Nevers, Y.; Rossier, V.; Train, C.; Altenhoff, A. M.; Dessimoz, C.; Glover, N.
Show abstract
Assessing the quality of protein-coding gene repertoires is critical in an era of increasingly abundant genome sequences for a diversity of species. State-of-the-art genome annotation assessment tools measure the completeness of a gene repertoire, but are blind to other types of errors, such as gene over-prediction or contamination. We developed OMArk, a software relying on fast, alignment-free sequence comparisons between a query proteome and precomputed gene families across the tree of life. OMArk assesses not only the completeness, but also the consistency of the gene repertoire as a whole relative to closely related species. It also reports likely contamination events. We validated OMArk with simulated data, then performed an analysis of the 1805 UniProt Eukaryotic Reference Proteomes, illustrating its usefulness for comparing and prioritizing proteomes based on their quality measures. In particular, we found strong evidence of contamination in 59 proteomes, and identified error propagation in avian gene annotation resulting from the use of a fragmented zebra finch proteome as reference. OMArk is available on GitHub (https://github.com/DessimozLab/OMArk), as a Python package on PyPi, and as an interactive online tool at https://omark.omabrowser.org/.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SHEPHARD: a modular and extensible software architecture for analyzing and annotating large protein datasets 94%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 93%
- Proteome-level assessment of origin, prevalence and function of Leucine-Aspartic Acid (LD) motifs 92%
Similar papers in this journal
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 94%
- The structural coverage of the human proteome before and after AlphaFold 93%
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 93%
Similar papers in this journal
- Purging genomes of contamination eliminates systematic bias from evolutionary analyses of ancestral genomes 95%
- Proteome allocation is linked to transcriptional regulation through a modularized transcriptome 93%
- PTMNavigator: Interactive Visualization of Differentially Regulated Post-Translational Modifications in Cellular Signaling Pathways 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.