DIAMOND2GO: A rapid Gene Ontology assignment and enrichment tool for functional genomics
Golden, C.; Studholme, D. J.; Farrer, R. A.
Show abstract
DIAMOND2GO (D2GO) is a new toolset to rapidly assign Gene Ontology (GO) terms to genes or proteins based on sequence similarity searches. D2GO uses DIAMOND for alignment, which is 100 - 20,000 X faster than BLAST. D2GO leverages GO- terms already assigned to sequences in the NCBI non-redundant database to achieve rapid GO-term assignment on large sets of query sequences. In one test, 98% of the 130,184 predicted human proteins and splice variants were assigned GO-terms (>2 million in total) in < 13 minutes on a laptop computer. D2GO also features the ability to perform enrichment analysis between subsets of data, thereby allowing rapid assignment and detection of over-represented GO-terms in novel sets of sequences. D2GO is freely available under the MIT licence from https://github.com/rhysf/DIAMOND2GO
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Flame (v2.0): advanced integration and interpretation of functional enrichment results from multiple sources 96%
- GOThresher: a program to remove annotation biases from protein function annotation datasets 95%
- MutaFrame - an interpretative visualization framework for deleteriousness prediction of missense variants in the human exome 94%
Similar papers in this journal
- Reciprocal Best Structure Hits: Using AlphaFold models to discover distant homologues 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 94%
- scExplorer: A Comprehensive Web Server for Single-Cell RNA Sequencing Data Analysis 93%
Similar papers in this journal
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 96%
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 95%
- Comprehensive benchmark of differential transcript usage analysis for static and dynamic conditions 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.