Back

Identifying intervention strategies from machine learning models with COALA: a counterfactual optimization framework

Han, B.; Duan, Q.; Hu, T.

2025-07-18 bioinformatics
10.1101/2025.07.18.664723 bioRxiv
Show abstract

MotivationMachine learning models in biomedicine have become increasingly complex, often functioning as black boxes. However, understanding contributors to disease and making actionable health interventions requires interpretable models. Common explainable AI methods like SHAP focus on feature importance but fall short in explaining why features contribute in certain patterns or what interventions to take. Counterfactual explanations address this by proposing "what if" scenarios but current tools focus on individual predictions and fail to generalize complex trends. ResultsWe propose the framework Counterfactual Optimization for Actionable interpretabiLity in AI (COALA). COALA interprets models by identifying optimal counterfactuals across user-defined mutable feature subsets and constraining remaining features to reveal how constraint features determine what interventions are optimal. By analyzing counterfactual profiles of features rather than individual features, COALA reveals holistic patterns. Using synthetic and real datasets, COALA reveals simple and complex model trends and provides more intuitive, multi-feature interventions than SHAP. Availability and ImplementationCode for COALA implementation, synthetic data, models trained on synthetic data, and code to replicate results and figures are available at https://github.com/brt-solo/COALA.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.