Back

Comparing the performance of functional versus taxonomic metagenomics for detecting ammonia disturbances in the biogas system

Boers, D.; Chapleur, O.; Andersson, A. F.; Schnürer, A.

2025-06-10 microbiology
10.1101/2025.06.10.658511 bioRxiv
Show abstract

Biogas is a renewable energy source with great potential, but its production is frequently hindered by process disturbances, of which a high ammonia concentration is one common cause. It is desirable that such disturbances are found as early as possible, and metagenomics data has the potential to improve this detection. This study compares functional and taxonomic aspects of metagenomics data, hypothesizing that functional data will perform better for detecting ammonia disturbances. The hypothesis was tested by metagenomic sequencing of samples from three independent studies, which followed lab-scale reactors during ammonia disturbances. The resulting sequences were used to predict protein-coding genes, which were functionally and taxonomically annotated. The read counts of these features were fitted to disturbance states and ammonia concentrations of reactor samples using regularized regression, which allowed filtering out irrelevant features even when the number of features was much larger than the number of samples. Taxonomic data had similar or better performance in detecting ammonia disturbances and in fitting ammonia concentrations, both when analyzing separate studies as well as when analyzing the combined data of the studies. Our hypothesis that functional metagenomics would outperform taxonomic metagenomics was therefore not supported. Visual abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=120 SRC="FIGDIR/small/658511v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@43aaeforg.highwire.dtl.DTLVardef@8b5d27org.highwire.dtl.DTLVardef@190e737org.highwire.dtl.DTLVardef@3c0206_HPS_FORMAT_FIGEXP M_FIG C_FIG

Published in FEMS Microbiology Ecology (predicted rank #16) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.