Back

How can spectrum bias impact expectations of test performance: a secondary modeling study estimating novel swab-based tuberculosis test outcomes across populations

Zaidi, S.; Garg, T.; Vo, L. N. Q.; Sander, M.; Codlin, A. J.; Byrne, R. L.; Lem, V.; Banu, S.; Bimba, J.; Santos, V. S.; Mbuli, C.; Nguyen, H. T.; Choko, A.; Wandiga, S.; John, S.; Squire, B.; Wingfield, T.; Creswell, J.

2026-07-04 public and global health
10.64898/2026.07.02.26357020 medRxiv
Show abstract

Background New swab-based near-point-of-care (NPOC) tests offer a potentially lower-cost, simpler alternative to Xpert MTB/RIF Ultra (Xpert-Ultra) testing for tuberculosis (TB) diagnosis. However, it is critical to understand how their performance may differ across diverse populations to inform programmatic rollout and real-world clinical decision-making. Methods We modeled the performance of testing sputum swabs with the MiniDock MTB assay and the Xpert MTB/RIF (Xpert) using positive percentage agreement (PPA) with Xpert-Ultra, disaggregated by semi-quantitative grade. PPA for MiniDock MTB was derived from a random-effects meta-analysis of three diagnostic accuracy studies. PPA for Xpert was derived from data from an early diagnostic study. PPA estimates were applied to 1,248 positive Xpert-Ultra test results from Phase 1 of the Start4All study across seven countries, disaggregated by facility-based (n=1,033) and community-based (n=215) participant recruitment, comparing a single overall PPA to an Xpert Ultra semi-quantitative grade-stratified model (from Trace to High). Results Pooled overall PPA was 84.8% (95% CI: 67.1-93.8%) for MiniDock MTB and 93.1% (90.3-95.2%) for Xpert. Grade-stratified modeling revealed lower PPAs at "Very Low" and "Trace" semi-quantitative grades: MiniDock MTB 56.5% and 33.8%; Xpert 66.7% and 23.1%, respectively. When grade-stratified estimates were applied to the Start4All data, both tests performed similarly, missing 241/1,248 and 228/1,248 positive Xpert-Ultra results, respectively. The single-value model overestimated performance most markedly in community settings, predicting 10.8% and 19.1% more positive results than the grade-stratified model for MiniDock MTB and Xpert, respectively. Conclusion MiniDock MTB sputum swabs perform comparably to Xpert when assessed using grade-stratified modeling. Both tests are likely to miss a greater proportion of people with Xpert-Ultra positive results in community settings, where paucibacillary disease is more common. Using single point estimates for diagnostic accuracy can substantially overestimate real-world performance, highlighting the importance of evaluating diagnostics across the full spectrum of TB disease.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.