Back

Discernibility in explanations: an approach to designing more acceptable and meaningful machine learning models for medicine

Wang, H.; Aligon, J.; May, J.; Doumard, E.; Labroche, N.; Delpierre, C.; Soule-Dupuy, C.; Casteilla, L.; Planat, V.; Monsarrat, P.

2025-04-10 bioinformatics
10.1101/2025.04.06.647498 bioRxiv
Show abstract

BackgroundAlthough the benefits of machine learning (ML) are undeniable in health-care, explainability plays a vital role in improving transparency and understanding the most decisive and persuasive variables for prediction. The challenge is to identify explanations that make sense to the biomedical expert. This work proposes discernibility as a new approach to faithfully reflect human cognition, with the users perception of a relationship between explanations and data for a given variable. MethodsA total of 50 participants (19 biomedical and 31 data scientists) evaluated their perception of the discernibility of explanations from both synthetic and human-based dataset (National Health and Nutrition Examination Survey). The inter-rater reliability was tested through the intraclass correlation coefficient (ICC). 13 statistical coefficients were considered to be able to capture for a given variable the relationship between its values and its explanations. A Passing-Bablok regression was performed for each user to highlight the consistency between user rating and each coefficient. FindingsThe low inter-rater reliability of discernibility (ICC{inverted exclamation}0.5) with no difference between areas of expertise or level of education underlines the need for an objective metric of discernibility. Among all evaluated metrics, dcor metric was found to be the most suitable to capture the intra-individual reliability of discernibility perceived by users (median slope closer to 1 and a narrower confidence interval width for the Passing-Bablok regression with the lowest differential bias between the most and least discernible values). Interpretationdcor was shown to be a reliable metric for assessing the discernibility of explanations, effectively capturing the clarity of the relationship between the data and their explanations, and providing clues to the underlying pathophysiological mechanisms that are not immediately apparent when examining individual predictors. Discernibility can also serve as an evaluation metric for model quality, used to prevent overfitting or aid in feature selection, providing medical practitioners with more accurate and persuasive results.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.