Back

Cohort-HMM marker recruitment with per-OG orthology QC for phylogenomic supermatrices

Nielsen, T. N.

2026-05-31 bioinformatics
10.64898/2026.05.27.728348 bioRxiv
Show abstract

OrthoFinders all-vs-all DIAMOND step systematically misses single-copy orthogroups (SC OGs) at deep taxonomic divergence: a marker recovered cleanly within a tightly defined cohort is dropped when the same marker is searched against phylum-broad metagenome-assembled genome (MAG) sets, because pairwise sequence similarity falls below DIAMONDs detection threshold even when the underlying ortholog is present. The result is biased dropout -- supermatrices that retain genomes near the cohort but lose genomes from the deeper, more diverged corners of the same phylum. We describe a two-stage cohort-HMM recruitment pipeline (per-OG profile HMMs built from cohort alignments, then hmmsearch against the broader proteome set) followed by an independent per-OG gene-tree QC step that classifies each recruited hit relative to the cohorts most recent common ancestor (MRCA) descendant set, with a per-MAG paralog-rate filter applied before supermatrix concatenation. We characterize the pipeline across three taxonomic ranks. At phylum scale (Omnitrophota, 97 cohort OGs, 714 NCBI MAGs), the recruitment recovers MAGs that the OrthoFinder-only supermatrix would otherwise drop, and the QC identifies 2 deep-peripheral MAGs -- divergent genomes whose per-OG tips repeatedly place outside the cohort MRCA descendant set despite being orthologs -- that the per-MAG filter removes. At family scale (Pelagibacteraceae, 146 cohort OGs, 366 NCBI MAGs) and at genus scale (Actinomarina, 289 cohort OGs, 23 NCBI MAGs), the per-tip paralog-candidate rate drops to 0.0 %. The pipeline addresses two independent failure modes. Cohort paralog density breaks strict-SC OG discovery at the cohort step (the family-rank case, where every candidate marker has at least one cohort species carrying multiple copies; the relaxed cohort criterion supplies the marker set and HMM recruitment disambiguates which copy each NCBI MAG contributes). DIAMOND-reach attrition breaks OG assignment for the most divergent NCBI MAGs (the phylum-rank case, where pairwise similarities fall below DIAMONDs detection threshold; HMM recruitment recovers the dropouts and the per-OG QC step filters residual paralog candidates). At genus rank both modes are inactive and OrthoFinder suffices directly; HMM recruitment runs but finds no new orthologs. Code and per-case data products are released as a community resource at Zenodo (DOI 10.5281/zenodo.20422348).

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nature Biotechnology
172 papers in training set
Top 0.1%
18.0%
2
Nature Methods
385 papers in training set
Top 0.5%
18.0%
3
Genome Biology
637 papers in training set
Top 1%
9.4%
4
Genome Research
468 papers in training set
Top 0.7%
7.1%
50% of probability mass above
5
Nucleic Acids Research
1281 papers in training set
Top 4%
5.3%
6
Nature Communications
5641 papers in training set
Top 31%
4.2%
7
Bioinformatics
1204 papers in training set
Top 5%
4.2%
8
Nature
645 papers in training set
Top 4%
3.9%
9
Molecular Biology and Evolution
542 papers in training set
Top 2%
3.1%
10
PLOS Computational Biology
1863 papers in training set
Top 11%
3.1%
11
Cell Systems
201 papers in training set
Top 2%
2.7%
12
Nature Genetics
286 papers in training set
Top 3%
1.9%
13
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.5%
14
Nature Microbiology
155 papers in training set
Top 2%
1.5%
15
Science
477 papers in training set
Top 6%
1.5%
16
Bioinformatics Advances
203 papers in training set
Top 4%
0.9%
17
PLOS ONE
5266 papers in training set
Top 63%
0.8%
18
BMC Bioinformatics
457 papers in training set
Top 6%
0.6%
19
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%
20
Cell
431 papers in training set
Top 12%
0.6%
21
Systematic Biology
144 papers in training set
Top 0.8%
0.6%
22
Nature Computational Science
55 papers in training set
Top 2%
0.6%
23
eLife
5828 papers in training set
Top 70%
0.6%