Back

A Reinforcement Learning Approach for Modeling Organic Compound-Induced Antimicrobial Resistance Dynamics

Hedman, H. D.

2026-01-20 microbiology
10.64898/2025.12.18.695270 bioRxiv
Show abstract

This study presented an exploratory reinforcement learning (RL)-based simulation framework for examining antimicrobial resistance (AMR) dynamics under repeated exposure to an organic antimicrobial stressor, using copper as a representative model compound. Within a simplified and explicitly constrained simulation environment, three agent strategies were evaluated: random action selection, a rule-based heuristic, and a tabular Q-learning agent. Simulations were conducted over fixed-length 40-cycle episodes in which agents adjusted copper exposure in response to evolving resistance-related state variables. Across experimental runs, the Q-learning agent exhibited lower cumulative antibiotic resistance burden, measured by the area under the curve (AUC) of minimum inhibitory concentration (MIC) values for chloramphenicol and polymyxin B, while also maintaining lower cumulative copper exposure relative to the rule-based and random baselines. The rule-based agent demonstrated intermediate performance, whereas the random agent showed higher variability and less stable resistance trajectories. These differences reflected divergence in simulated resistance dynamics over time rather than short-term fluctuations in resistance burden. Rather than providing predictive or mechanistic insight into microbial evolution, this work introduced an interpretable RL-based simulation framework intended to support comparative evaluation of sequential decision-making strategies under constrained observability, where high-resolution diagnostics or detailed biological measurements may be unavailable. Together, the results supported the use of reinforcement learning as a flexible methodological framework for studying AMR dynamics as a feedback-driven control problem under simplified and transparent assumptions.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.