Back

DataSAIL: Data Splitting Against Information Leakage

Joeres, R.; Blumenthal, D. B.; Kalinina, O. V.

2023-11-17 bioinformatics
10.1101/2023.11.15.566305 bioRxiv
Show abstract

Information Leakage is an increasing problem in machine learning research. It is a common practice to report models with benchmarks, comparing them to the state-of-the-art performance on the test splits of datasets. If two or more dataset splits contain identical or highly similar samples, a model risks simply memorizing them, and hence, the true performance is overestimated, which is one form of Information Leakage. Depending on the application of the model, the challenge is to find splits that minimize the similarity between data points in any two splits. Frequently, after reducing the similarity between training and test sets, one sees a considerable drop in performance, which is a signal of removed Information Leakage. Recent work has shown that Information Leakage is an emerging problem in model performance assessment. This work presents DataSAIL, a tool for splitting biological datasets while minimizing Information Leakage in different settings. This is done by splitting the dataset such that the total similarity of any two samples in different splits is minimized. To this end, we formulate data splitting as a Binary Linear Program (BLP) following the rules of Disciplined Quasi-Convex Programming (DQCP) and optimize a solution. DataSAIL can split one-dimensional data, e.g., for property prediction, and two-dimensional data, e.g., data organized as a matrix of binding affinities between two sets of molecules, accounting for similarities along each dimension and missing values. We compute splits of the MoleculeNet benchmarks using DeepChem, the LoHi splitter, GraphPart, and DataSAIL to compare their computational speed and quality. We show that DataSAIL can impose more complex learning tasks on machine learning models and allows for a better assessment of how well the model generalizes beyond the data presented during training.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.