Similar performance of 8 machine learning models on 71 censored medical datasets: a case for simplicity
Rebaud, L.; Capobianco, N.; Captier, N.; Escobar, T.; Spottiswoode, B.; Buvat, I.
Show abstract
In the analysis of medical data with censored outcomes, identifying the optimal machine learning pipeline is a challenging task, often requiring extensive preprocessing, feature selection, model testing, and tuning. To investigate the impact of the choice of pipeline on prediction performance, we evaluated 9 machine learning models on 71 medical datasets with censored targets. Only the decision tree model was consistently underperforming, while the other 8 models performed similarly across datasets, with little to no improvement from preprocessing optimization and hyperparameter tuning. Interestingly, more complex models did not outperform simpler ones, and reciprocally. ICARE, a straightforward model univariately learning only the sign of each feature instead of a weight, demonstrated similar performance to other models across most datasets while exhibiting lower overfitting, particularly in high-dimensional datasets. These findings suggest that using the ICARE model to build signatures between centers could improve reproducibility. Our findings also challenge the traditional approach of extensive model testing and tuning to improve performance.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accurate Prediction of Breast Cancer Survival through Coherent Voting Networks with Gene Expression Profiling 95%
- Synthetic data for privacy-preserving clinical risk prediction 94%
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 94%
Similar papers in this journal
- Data-driven Discovery of Mathematical and Physical Relations in Oncology Data using Human-understandable Machine Learning 93%
- Survival Prediction Landscape: An In-Depth Systematic Literature Review on Activities, Methods, Tools, Diseases, and Databases 93%
- An Explainable Multi-Modal Neural Network Architecture for Predicting Epilepsy Comorbidities Based on Administrative Claims Data 91%
Similar papers in this journal
- Towards Predicting 30-Day Readmission among Oncology Patients: Identifying Timely and Actionable Risk Factors 92%
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 92%
- A Bayesian Framework for Detecting Gene Expression Outliers in Individual Samples 92%
Similar papers in this journal
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 95%
- On the predictability of postoperative complications for cancer patients: a Portuguese cohort study 93%
- Combining symbolic regression with the Cox proportional hazards model improves prediction of heart failure deaths 93%
Similar papers in this journal
- Reliable machine learning models in genomic medicine using conformal prediction 94%
- Improved Predictions Of MHC-Peptide Binding Using Protein Language Models 92%
- BC-Predict: Mining of signal biomarkers and multilevel validation of cascade classifier for early-stage breast cancer subtyping and prognosis 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.