Back

From Real-World Data to Virtual Intervention: A Probabilistic Neural Network for Simulating Kidney Function Preservation via Proteinuria Reduction

Takeda, A.; Igata, H.; Mizuno, K.; Yano, Y.; Nagasu, H.; Ohashi, M.; Kashihara, N.; Kobayashi, H.

2026-07-15 nephrology
10.64898/2026.07.12.26357786 medRxiv
Show abstract

Predicting the long-term kidney function decline is critical for timely intervention but remains challenging. While the urinary protein-to-creatinine ratio (uPCR) is a potential surrogate endpoint, its short-term reduction's link to long-term nephroprotection requires investigation. This study aimed to develop a probabilistic neural network model to capture both the estimated glomerular filtration rate (eGFR) slope and its uncertainty based on baseline clinical characteristics. Using a retrospective dataset, we designed a neural network to output a predictive distribution (mean and standard deviation {sigma}) for the eGFR slope. SHAP (SHapley Additive exPlanations) was used for model interpretation, and a simulation study quantified the impact of uPCR reduction. In the validation set, the model achieved a Pearson's correlation coefficient of 0.56 and an RMSE of 2.81 ml/min/1.73m^2/year between predicted and actual slopes. SHAP analysis identified uPCR as the most potent predictor, with higher baseline levels associated with a more rapid eGFR decline. Furthermore, a simulated 62% uPCR reduction demonstrated a significant improvement in the predicted eGFR slope, an effect most pronounced in patients with high baseline uPCR. This proof-of-concept study reinforces the critical role of uPCR in predicting eGFR slope and suggests its reduction may contribute to long-term kidney function preservation, warranting validation in larger, diverse real-world datasets.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.