Deviation Error: assessing machine learning predictions for replicate measurements in genomics and beyond
Abdulnabi, H.; Westwood, J. T.
Show abstract
A quantitative measurement can have variation, referred to here as measurement variation, which is a probability distribution. Machine Learning models typically produce a prediction corresponding to the mode of the measurement variation. The Deviation Error is a novel metric, described here, to assess predictions that accounts for measurement variation. Measurement variations in genomics data were explored. Towards a general prescription for modelling genomics measurements, different loss functions were used to fit models on synthetically generated data that mimics genomics measurements. Synthetically generated data offers the ability to know the true underlying value and to control the forms and amounts of noise injected at different stages of data processing. Different datasets were generated with varying levels of noise. Of the loss functions tried, only models fit with the Deviation Error performed as well if not better on any of the combinations of the metrics and datasets.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.