DataSAIL: Data Splitting Against Information Leakage
Joeres, R.; Blumenthal, D. B.; Kalinina, O. V.
Show abstract
Information Leakage is an increasing problem in machine learning research. It is a common practice to report models with benchmarks, comparing them to the state-of-the-art performance on the test splits of datasets. If two or more dataset splits contain identical or highly similar samples, a model risks simply memorizing them, and hence, the true performance is overestimated, which is one form of Information Leakage. Depending on the application of the model, the challenge is to find splits that minimize the similarity between data points in any two splits. Frequently, after reducing the similarity between training and test sets, one sees a considerable drop in performance, which is a signal of removed Information Leakage. Recent work has shown that Information Leakage is an emerging problem in model performance assessment. This work presents DataSAIL, a tool for splitting biological datasets while minimizing Information Leakage in different settings. This is done by splitting the dataset such that the total similarity of any two samples in different splits is minimized. To this end, we formulate data splitting as a Binary Linear Program (BLP) following the rules of Disciplined Quasi-Convex Programming (DQCP) and optimize a solution. DataSAIL can split one-dimensional data, e.g., for property prediction, and two-dimensional data, e.g., data organized as a matrix of binding affinities between two sets of molecules, accounting for similarities along each dimension and missing values. We compute splits of the MoleculeNet benchmarks using DeepChem, the LoHi splitter, GraphPart, and DataSAIL to compare their computational speed and quality. We show that DataSAIL can impose more complex learning tasks on machine learning models and allows for a better assessment of how well the model generalizes beyond the data presented during training.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian Networks 97%
- Predicting Affinity Through Homology (PATH): Interpretable Binding Affinity Prediction with Persistent Homology 97%
- Interpretable Pairwise Distillations for Generative Protein Sequence Models 96%
Similar papers in this journal
- Building explainable graph neural network by sparse learning for the drug-protein binding prediction 96%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 96%
- Representation of k-mer sets using spectrum-preserving string sets 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.