Back

GenomeCompendium: A database for the integrated analysis of repeats, assembly quality and functional content of complete prokaryotic genomes

Totu, T.; Jaques, G.; Heiniger, B.; Segessemann, T.; Schmid, M.; Bourqui, M.; Wicki, A.; Frey, J. E.; Ahrens, C. H.

2026-08-19 genomics
10.64898/2026.08.14.744864 bioRxiv
Show abstract

Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes (https://genome-compendium.com/). Using NCBI RefSeq ([~]47,000) and GenBank ([~]13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classification and frequency analysis, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links [~]90 features across genomes, offers downloadable reports and -as unique features-pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.