Back

Orthology transfer maps only the conserved core of the Varroa destructor proteome and over-calls host absence two times in three

Ryba, S.

2026-08-26 bioinformatics
10.64898/2026.08.20.745999 bioRxiv
Show abstract

The ectoparasitic mite Varroa destructor is the principal threat to managed honey bees, and a test case for the genome-scale methods applied to non-model organisms, nearly all of which infer from orthology. We reconstructed the first genome-wide protein-interaction network for V. destructor (7,080 proteins, 335,914 interactions), whose modular structure exceeds a degree-preserving null by 368 standard deviations, but whose every edge is interolog-transferred and every node conserved at least to Eukaryota. None of the 791 genes lacking an orthologous group enters it - arithmetic rather than discovery - yet the excluded compartment is large and coherent. It comprises 3,161 genes (30.9% of the proteome), shorter and less annotated than the rest; an annotation-free genome search detects orphans in a tick genome at 4.0% against 70.4% for networked genes. Within the orthology-bearing compartment visibility is non monotonic: the Acari-level bin (74.1%) falls below the Arthropoda-level bin (89.9%). The same logic applied to host comparison yields a benchmarked error: of genes called absent from Apis on group identity alone, 67.4% recover a sequence homologue - against zero for a shuffled null and 1.3% in the presence direction - rising to 78.3% in the least panel-biased stratum. Both figures are properties of the calling rule: under an identity floor the error directions cross near 34% identity; orthology cannot be said to err in either direction without fixing the criterion first. Host divergence resolves into gene absence and residue level substitution, falling in those two compartments respectively. A bee-sparing target map follows as broader impact.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.