Back

OGGfinder: Accurate Orthogroup Inference for Pan-Gene Families in Complex Genomes

LIU, F.; Wang, H.

2026-06-11 genomics
10.64898/2026.06.08.726684 bioRxiv
Show abstract

Accurate inference of orthologous gene groups (OGGs) is a foundational step in comparative genomics, yet existing tools fail to meet the demands of complex allopolyploid genomes. Here, we present OGGfinder, a novel pipeline that integrates sequence similarity with phylogenetic tree topology constraints, a data-driven 5th percentile (P5) threshold inference mechanism, a robust 6-step post-processing pipeline, and Latin Hypercube Sampling (LHS) for automated parameter optimization. In a benchmark utilizing 2,920 AP2 gene family from 164 allopolyploid cotton (Gossypium) genomes with a target orthogroup size of 164 genes, OGGfinder successfully recovered 18 high-quality OGGs with a mean size of 162.2 genes and zero singletons, tightly approximating the expected species count. In contrast, OrthoFinder drastically over-clustered genes into only 7 massive groups (mean size 417.1, max 654), while TreeCluster heavily fragmented the data into 116 groups with 42 singletons (36.2% singleton rate). CD-HIT generated 22 groups with a median size of only 52.5 and discarded nearly 1,000 sequences (34.6% gene loss) due to its greedy redundancy-reduction strategy. Comprehensive six-dimensional evaluation (completeness, granularity, topology consistency, auto-parameterization, polyploidy support, and scalability) yielded total scores of OGGfinder 26.0, OrthoFinder 23.9, CD-HIT 20.0, and TreeCluster 13.7 out of 30. These results demonstrate that OGGfinder significantly outperforms existing state-of-the-art tools, offering a highly accurate and reproducible solution for pan-gene family analyses in polyploid species. This is particularly critical for the application of finding OGGs within pan-gene families.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.