FIDDL: depth-matched negative controls distinguish genuine interspecific introgression from competitive-mapping artifact
Taylor, K.; Shumaker, K. A.; Gray, S. J.; Bochman, M. L.
Show abstract
Interspecific introgression is routinely detected by competitively mapping reads to a concatenated multi-species reference and calling regions where a non-focal species recruits coverage. Using strains that cannot contain the ancestry being detected, we show this design generates substantial false-positive signal through two mechanisms with opposite phylogenetic-distance signatures. Standard nuclear assemblies omit the mitochondrion and 2-micron plasmid, leaving high-copy cytoplasmic reads without a legitimate target; completing the reference preferentially removes signal from the most divergent donor. Genuine cross-species sequence conservation inflates the most closely related donor. Masking chromosome ends removes its subtelomeric part but plateaus at a non-zero floor, and the interior residual traces to conserved single-copy genes where a short read carries under one base of discriminating information. The floor grows with sequencing depth (1.19% of callable positions at 50x, 2.02% at 147x, 3.85% at 393x in a pure strain), is not mitigated by long reads, and appears at sub-diploid dosage - three properties widely read as evidence of authenticity. Because the discriminating information is below single-read resolution, no read-level filter separates artifact from introgression; we show three that fail. What works is locus-level: a consensus-phylogenetic test (29/29 specificity on confirmed artifact) and an allele-fraction donor-match test, complementary and validated in both directions on independent published introgression. We package the comparative controls as FIDDL (False Introgression Detection via Depth-matched controls and Loci-recurrence), an open-source tool, withdraw two of our own analysis-ready calls, and show re-analysis of published wild isolates reduces low-confidence introgression by [~]53% while leaving high-confidence signal intact.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.