A Clear, Legible, Explainable, Transparent, and Elucidative (CLETE) Binary Classification Platform for Tabular Data
Nasimian, A.; Younus, S.; Hammarlund, E. U.; Pienta, K. J.; Rönnstrand, L.; Kazi, J. U.
Show abstract
Therapeutic resistance continues to impede overall survival rates for those affected by cancer. Although driver genes are associated with diverse cancer types, a scarcity of instrumental methods for predicting therapy response or resistance persists. Therefore, the impetus for designing predictive tools for therapeutic response is crucial and tools based on machine learning open new opportunities. Here, we present an easily accessible platform dedicated to Clear, Legible, Explainable, Transparent, and Elucidative (CLETE) yet wholly modifiable binary classification models. Our platform encompasses both unsupervised and supervised feature selection options, hyperparameter search methodologies, under-sampling and over-sampling methods, and normalization methods, along with fifteen machine learning algorithms. The platform furnishes a k-fold receiver operating curve (ROC) - area under the curve (AUC) and accuracy plots, permutation feature importance, SHapley Additive exPlanations (SHAP) plots, and Local Interpretable Model-agnostic Explanations (LIME) plots to interpret the model and individual predictions. We have deployed a unique custom metric for hyperparameter search, which considers both training and validation scores, thus ensuring a check on under or over-fitting. Moreover, we introduce an innovative scoring method, NegLog2RMSL, which incorporates both training and test scores for model evaluation that facilitates the evaluation of models via multiple parameters. In a bid to simplify the user interface, we provide a graphical interface that sidesteps programming expertise and is compatible with both Windows and Mac OS. Platform robustness has been validated using pharmacogenomic data for 23 drugs across four diseases and holds the potential for utilization with any form of tabular data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 95%
- Drug Response Prediction and Biomarker Discovery Using Multi-Modal Deep Learning 94%
- Identification of Monotonically Classifying Pairs of Genes for Ordinal Disease Outcomes 94%
Similar papers in this journal
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 94%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 94%
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 94%
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 95%
- Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning algorithms for patient classification using metabolomics biomedical data. 95%
- Network-based estimation of therapeutic efficacy and adverse reaction potential for prioritisation of anti-cancer drug combinations 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.