Back

Predicting Antifouling Paint Particle Contamination based on 16S rRNA Gene Sequencing Data using Random Forest-Based Machine Learning

Tagg, A. S.; Sperlea, T.; Labrenz, M.; Schenk, A.; Kreikemeyer, B.

2025-09-03 microbiology
10.1101/2025.09.03.673959 bioRxiv
Show abstract

Antifouling paints often contain biocides designed to inhibit biological growth, and antifouling paint particles (APPs) have been previously shown to affect microbial communities in sediment. Given typical methods for monitoring for APP presence can be specialised and challenging, alternative methods using simple, standardised and universal approaches, such as 16S rRNA amplicon sequencing, would be highly valuable. This study uses a field-based mesocosm approach to train a random forest-based (supervised) machine learning model to predict APPs presence and concentration in sediment based on 16S microbial community data. The model correctly predicted 100% APP-presence samples and 83.3% APP-absence samples in the incubation testset, although the model could not correctly predict APP concentration with sufficient accuracy. To determine real-world applicability of the model, samples from 14 different sites along the Baltic Sea coastline and Warnow estuary in NE Germany were collected and APP-presence was pre-determined using SEM-EDX spectroscopy. The model correctly assigned APP-absence status to all APP-absent sites, and correctly assigned 3 of 5 APP-contaminated sites as having APP presence. As such, these results confirm it is possible to predict APPs in sediment based on the microbial community, and serves as a proof-of-concept for the further development of machine learning-based predictive tools for environmental monitoring.

Published in Microbiology Spectrum (predicted rank #13) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.