Back

A calibrated novelty flag for fungal ITS metabarcoding: choosing the error rate at which sequences are declared new

O'Brien, A.; Parada, P.

2026-08-04 bioinformatics
10.64898/2026.07.29.741524 bioRxiv
Show abstract

O_LIEnvironmental fungal surveys routinely recover internal transcribed spacer (ITS) sequences that cannot be assigned at fine taxonomic ranks, the so-called fungal "dark matter." Such sequences are set aside by thresholding a similarity or confidence score at a conventional value. Those conventions do abstain, but the error rate a threshold implies is neither stated nor selectable, and a threshold defined for one kind of score does not transfer to another. C_LIO_LIWe present a conformal novelty flag that supplies what is missing: each query receives a p-value with a distribution-free guarantee that the rate of falsely declaring a known sequence novel is bounded by a user-chosen . We evaluate it on a leave-one-genus-out benchmark built from UNITE, on alignment identity, a k-mer bootstrap consensus and two neural classifiers output probabilities, and on a soil fungal dataset. C_LIO_LIThe flag holds its nominal rate across two orders of magnitude in , so an operating point can be chosen rather than inherited: at = 0.05 it fires on 4.3% of known-genus queries and recovers 19.2% of genuinely novel genera. The cutoff holding a 5% error rate here is 64.8% identity, nowhere near the customary 97%, showing how little a threshold carries its error rate between datasets. Coverage transferred across eleven settings spanning those four scores, two amplicon regions and a sevenfold change in reference size, all within 1.1 percentage points of nominal, while detection ranged from 5.2% to 53.3%: the guarantee is on the error rate and not on power, and two of our settings are valid but uninformative. Applied to soil data the flag identifies 20.7% of amplicon sequence variants as novel at a controlled 5% error rate, 16.3% under an abundance filter. Half of those recur near-identically among GlobalFungis unnamed environmental variants while fewer than one in ten matches a named species hypothesis, a sixfold skew towards the uncatalogued against 1.9-fold for sequences the flag passes. C_LIO_LIThe flag turns an arbitrary cutoff into a decision with a stated error rate, and in doing so converts dark matter from a residue into a set of prioritizable targets for formal description. C_LI

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.