Quantification of individual dataset contributions to prediction accuracy in cooperative learning
Dechantsreiter, D.; Kelly, R.; Lange, C.; Lasky-Su, J. A.; Hahn, G.
Show abstract
We consider cooperative learning, a recently proposed technique to leverage the predictive power of several datasets (also called data views) to improve the prediction accuracy of an outcome of interest. Cooperative learning uses a Lasso-type penalty to fit several datasets to a given outcome, while a so-called agreement penalty enforces that the predictions made by the individual datasets agree. We are interested in the question of whether the predictive power of each individual dataset can be quantified using cooperative learning. In this work, we demonstrate that certain trace plots, analogously to the ones for the classic Lasso, allow one to quantify the predictive power of each individual dataset. Importantly, this allows one to detect datasets which do not carry any predictive power on the outcome. In an experimental study, we quantify the predictive power of three real datasets in the context of the Childhood Asthma Management Program (CAMP), with the three datasets containing information on epidemiological variables, metabolites, and clinical data, respectively.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accessibility of covariance information creates vulnerability in Federated Learning frameworks 96%
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 95%
- LSMMD-MA: Scaling multimodal data integration for single-cell genomics data analysis 95%
Similar papers in this journal
- Towards the Genome-scale Discovery of Bivariate Monotonic Classifiers 95%
- A Comparison of Embedding Aggregation Strategies in Drug-Target Interaction Prediction 94%
- Transfer Learning Models for Bacterial Strain Dissemination Biomarkers using Weighted Non-Parallel Proximal Support Vector Machines 94%
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 95%
- Analyzing Biomarker Discovery: Estimating the Reproducibility of Biomarker Sets 94%
- Theoretical properties of nearest-neighbor distance distributions and novel metrics for high dimensional bioinformatics data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.