Back

Paired evaluation defines performance landscapes for machine learning models

Nariya, M. K.; Mills, C. E.; Sorger, P. K.; Sokolov, A.

2022-09-12 bioinformatics
10.1101/2022.09.07.507020 bioRxiv
Show abstract

The true accuracy of a machine learning model is a population-level statistic that cannot be observed directly. In practice, predictor performance is estimated against one or more test datasets, and the accuracy of this estimate strongly depends on how well the test sets represent all possible unseen datasets. Here we present paired evaluation, a simple approach for increasing the robustness of performance evaluation by systematic pairing of test samples, and use it to evaluate predictors of drug response in breast cancer cell lines and of disease severity in patients with Alzheimers Disease. Our results demonstrate that the choice of test data can cause estimates of performance to vary by as much as 30%, and that paired evaluation makes it possible to identify outliers, improve the accuracy of performance estimates in the presence of known confounders, and assign statistical significance when comparing machine learning models.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Cell Systems
201 papers in training set
Top 0.1%
34.0%
2
Scientific Reports
3612 papers in training set
Top 9%
7.2%
3
Nature Communications
5641 papers in training set
Top 22%
7.2%
4
PLOS Computational Biology
1863 papers in training set
Top 5%
6.7%
50% of probability mass above
5
Bioinformatics
1204 papers in training set
Top 5%
4.0%
6
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.4%
7
Patterns
78 papers in training set
Top 0.5%
3.4%
8
Communications Biology
993 papers in training set
Top 6%
3.1%
9
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 26%
1.9%
10
iScience
1154 papers in training set
Top 17%
1.7%
11
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
12
BMC Bioinformatics
457 papers in training set
Top 4%
1.5%
13
npj Breast Cancer
23 papers in training set
Top 0.3%
1.5%
14
Nature Methods
385 papers in training set
Top 4%
1.5%
15
eLife
5828 papers in training set
Top 52%
1.5%
16
Science Advances
1243 papers in training set
Top 22%
1.4%
17
Nature Machine Intelligence
70 papers in training set
Top 2%
1.4%
18
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.6%
1.1%
19
npj Systems Biology and Applications
125 papers in training set
Top 1%
1.1%
20
Human Genetics and Genomics Advances
84 papers in training set
Top 2%
1.0%
21
Cell Reports Methods
165 papers in training set
Top 3%
1.0%
22
npj Digital Medicine
118 papers in training set
Top 3%
0.9%
23
Cell Reports Medicine
153 papers in training set
Top 5%
0.8%
24
Molecular Systems Biology
162 papers in training set
Top 3%
0.8%
25
Cancer Cell
42 papers in training set
Top 2%
0.6%
26
Cancer Research
130 papers in training set
Top 3%
0.6%