Back

A comprehensive benchmark of discrepancies across microbial genome reference databases

Boldirev, G.; Aguma, P.; Munteanu, V.; Koslicki, D.; Alser, M.; Zelikovsky, A.; Mangul, S.

2026-03-04 bioinformatics
10.64898/2026.03.02.709103 bioRxiv
Show abstract

Metagenomic analysis of microbial communities relies significantly on the quality and completeness of reference genomes, which allow researchers to compare sequencing reads against reference genome collections to reveal essential community characteristics. However, the reliability of these analyses is often compromised by substantial discrepancies across existing reference resources, including differences in genome content, assembly fragmentation, taxonomic representation, and metadata completeness. While these inconsistencies are known to introduce bias, the extent of divergence between major databases remains largely unknown. Here, we present a comprehensive benchmark of discrepancies across multiple widely used microbial genome reference resources. We developed the Cross-DB Genomic Comparator (CDGC), which utilizes reference genome alignments to systematically capture discrepancies in genome assemblies across reference databases. Applying this framework, we found that 99% of viral genomes were identical across databases, indicating strong consistency in viral reference resources. In contrast, fungal genomes showed substantially greater variability: although 82% of assemblies exhibited at least 90% similarity, only 7% were identical across databases. More concerning, we identified a subset of 461 assemblies with less than 50% similarity, suggesting the presence of technical artifacts, incomplete assemblies, or damaged genome files that require closer examination. Collectively, these results demonstrate that systematic cross-database benchmarking provides a critical mechanism for refining the accuracy of individual reference databases and advancing efforts towards more unified and reliable universal reference genomes.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.