Back

An integrated analytical framework for gender-based violence research: A simulation study combining machine learning and causal inference

Mboya, G. O.

2025-12-18 epidemiology
10.64898/2025.12.15.25342247 medRxiv
Show abstract

BackgroundCurrent research on Gender-Based Violence (GBV) typically separates predictive machine learning and causal inference into distinct analytical silos. Yet, grasping the multi-level determinants of violence requires an approach that can both identify high-value predictors and disentangle their causal mechanisms. MethodsThis study develops and demonstrates an integrated five-phase analytical framework for GBV research, applying sequential methods to a single synthetic dataset (N = 3,000) parameterized to reflect the prevalence patterns and risk factor distributions of the 2022 Kenya Demographic and Health Survey (KDHS). The framework was applied across five stages: descriptive epidemiology, Random Forest variable selection, logistic regression for adjusted associations, mediation analysis, and evidence synthesis. ResultsThe Random Forest model achieved 74.6% accuracy (AUC = 0.711; sensitivity = 37.3%; specificity = 89.8%), recovering partner alcohol use, childhood trauma, and marital conflict as the top-ranked predictors. Importantly, this accuracy offers only a modest gain over the null classifier baseline of 73.1%. Multivariable logistic regression yielded stable adjusted effect estimates for partner alcohol use (aOR = 6.60) and childhood trauma (aOR = 1.99). Mediation analysis tested three theoretically informed indirect pathways; however, none of the hypothesized indirect effects reached statistical significance (ACME p > 0.05 for all pathways), indicating that the specified risk factors operated primarily through direct rather than mediated routes in this simulation. ConclusionsThis proof-of-concept demonstration shows that the five-phase framework can be coherently applied to synthetic GBV data parameterized from KDHS estimates. The frameworks multi-level logistic regression approach produced meaningfully better model fit than single-level alternatives ({Delta}AIC = 67.5). Future research should apply this framework to real-world longitudinal data to rigorously evaluate its performance characteristics and causal inference capacity.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.