miniBUSCO: a faster and more accurate reimplementation of BUSCO
Huang, N.; Li, H.
Show abstract
MotivationAssembly completeness evaluation of genome assembly is a critical assessment of the accuracy and reliability of genomic data. An incomplete assembly can lead to errors in gene predictions, annotation, and other downstream analyses. BUSCO is one of the most widely used tools for assessing the completeness of genome assembly by comparing the presence of a set of single-copy orthologs conserved across a wide range of taxa. However, the runtime of BUSCO can be long, particularly for some large genome assemblies. It is a challenge for researchers to quickly iterate the genome assemblies or analyze a large number of assemblies. ResultsHere, we present miniBUSCO, an efficient tool for assessing the completeness of genome assemblies. miniBUSCO utilizes the protein-to-genome aligner miniprot and the datasets of conserved orthologous genes from BUSCO. Our evaluation of the real human assembly indicates that miniBUSCO achieves a 14-fold speedup over BUSCO. Furthermore, miniBUSCO reports a more accurate completeness of 99.6% than BUSCOs completeness of 95.7%, which is in close agreement with the annotation completeness of 99.5% for T2T-CHM13. Availabilityhttps://github.com/huangnengCSU/minibusco. Contacthli@ds.dfci.harvard.edu Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 95%
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 95%
- MoGAAAP: A modular Snakemake workflow for automated genome assembly and annotation with quality assessment 94%
Similar papers in this journal
Similar papers in this journal
- SwiftOrtho: a Fast, Memory-Efficient, Multiple Genome Orthology Classifier 95%
- Sequence Compression Benchmark (SCB) database - a comprehensive evaluation of reference-free compressors for FASTA-formatted sequences 95%
- CoCoPyE: feature engineering for learning and prediction of genome quality indices 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.