Back

Circulating Microbial DNA as a Potential Cancer Biomarker: Technical Challenges and Controlled Evaluation

Fry Brumit, D.; Bsteh, D.; Sun, S.; Goodrich, J. A.; Alderete, T. L.; Liss, M. A.; Goldkorn, A.; Fodor, A. A.

2026-08-20 bioinformatics
10.64898/2026.08.17.745053 bioRxiv
Show abstract

Abstract Background. Circulating microbial DNA (cmDNA) has been proposed as a non-invasive cancer biomarker, but most evidence comes from cancer-sequencing datasets not designed for microbial analysis and lacking contamination controls. Whether reported signatures reflect biology or artifact is unclear in low-biomass specimens, where standard taxonomic pipelines are prone to systematic error. Methods. In a tightly controlled pilot study of metastatic castration-resistant prostate cancer, we profiled plasma cell-free DNA (cfDNA) and buffy-coat genomic DNA (gDNA) from two patients and two healthy volunteers alongside mock blood-draw and reagent controls, each with and without host-DNA depletion. Reads were classified with Kraken2/Bracken and, independently, with the marker-gene classifier MetaPhlAn. As informatics controls, reads were per-base shuffled to randomize nucleotide order while preserving read length and guanine-cytosine (GC) content, and purely synthetic reads were generated from a four-base process matched only to an aggregate GC target; both were classified identically. Genus abundances were regressed against Kraken2 database k-mer representation and against GC content. Results. Across 40 samples, Kraken2 reported several thousand genera, samples clustered by specimen type in principal-coordinate analysis (PCoA), and pooled genus counts correlated strongly with a published cancer-microbiome catalog (The Cancer Genome Atlas lung adenocarcinoma, TCGA-LUAD; Spearman {rho} = 0.81 over 282 shared genera), a pattern readily interpreted as biological signal. However, these observations were also made in per-base shuffling, which preserves GC content and length but destroys all biological sequence: shuffled reads were still abundantly classified, still clustered by specimen type, and still correlated with the catalog ({rho} {approx} 0.7), as did every sample group, including pure reagent controls. Genus counts scaled tightly with each genus's k-mer representation in the Kraken2 database on real (r 2 = 0.74) and shuffled (r2 = 0.85) reads, and the same dependence appeared in the independent published cohort. Purely synthetic reads carrying no information beyond an aggregate GC target reproduced much of the cross-cohort agreement (synthetic TCGA-LUAD {rho} = 0.61 versus 0.81 for real reads; significant in 27 of 33 TCGA cancers), and replicate shuffles of a low-GC versus a high-GC plasma sample, for which the true difference is zero, produced spurious significant differences in about 46% of genera. Regressing observed counts against the shuffled baseline left 23 genera above the artifact floor at 5% false discovery rate (FDR), nearly all known kit contaminants, control-enriched viruses, or very-low-abundance taxa; a four-criterion validity filter reduced thousands of Kraken2 genera to a single defensible candidate, Klebsiella. Conclusions. Much of the apparent cmDNA structure, including its agreement with a published cancer-microbiome catalog, is explained by base composition and reference-database architecture rather than authentic biology, and short-read k-mer pipelines cannot separate the two on their own. We find little positive evidence of an authentic circulating microbial signal, though our small sample cannot prove its absence. To limit false discovery in low-biomass metagenomics, we recommend specimen-matched negative controls, corroboration with a conservative second classifier, per-base shuffling (with GC-matched synthetic reads as a stricter floor), and GC-aware analysis.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.