Back

Agentic-TimesFM-AKI: A Dual LLM-Time Series Framework for Predicting Drug-Induced Acute Kidney Injury with Privacy-Preserving Synthetic Data

AL-Sakkaf, G. E.

2026-07-31 health informatics
10.64898/2026.07.30.26359271 medRxiv
Show abstract

Background: Acute kidney injury (AKI) is a severe complication in intensive care units, frequently exacerbated by synergistic nephrotoxicity from drugs such as Vancomycin and Piperacillin-Tazobactam. Traditional alert systems relying on static thresholds suffer from high false-positive rates and delayed detection. Methods: We developed Agentic-TimesFM-AKI, a dual-model architecture integrating a Large Language Model (Gemma-4 Sentinel) with a zero-shot time-series forecaster (TimesFM) to provide continuous, dynamic risk forecasting and transparent clinical reasoning. The system was trained on a synthetically generated cohort with differential privacy ({varepsilon}=10) and evaluated on the publicly accessible eICU (N=200) and MIMIC-IV (N=117) Demo datasets. Results: In the internal eICU pilot evaluation, the framework achieved an Accuracy of 0.970 (95% CI: 0.945-0.990) and an F1-Score of 0.966, successfully mapping temporal physiological trajectories into intelligible natural language alerts. However, external validation on the MIMIC-IV cohort revealed severe performance degradation. Conclusions: While the dual-model framework provides highly accurate and interpretable AKI alerts on familiar schema cohorts, it suffers from structural formatting fragility and domain shift. This highlights critical vulnerabilities in applying generative models to out-of-distribution electronic health records. Keywords: Acute Kidney Injury, Large Language Models, Time-Series Forecasting, Electronic Health Records, Differential Privacy, Pharmacovigilance.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.