Back

Training a neural network to learn other dimensionality reduction removes data size restrictions in bioinformatics and provides a new route to exploring data representations

Dexter, A.; Thomas, S. A.; Steven, R. T.; Robinson, K. N.; Taylor, A. J.; Elia, E.; Nikula, C.; Campbell, A. D.; Panina, Y.; Najumudeen, A. K.; Murta, T.; Yan, B.; Grabowski, P.; Hamm, G.; Swales, J.; Gilmore, I.; Yuneva, M.; Goodwin, R. J. A.; Barry, S.; Sansom, O. J.; Takats, Z.; Bunch, J.

2020-09-03 bioinformatics
10.1101/2020.09.03.269555 bioRxiv
Show abstract

High dimensionality omics and hyperspectral imaging datasets present difficult challenges for feature extraction and data mining due to huge numbers of features that cannot be simultaneously examined. The sample numbers and variables of these methods are constantly growing as new technologies are developed, and computational analysis needs to evolve to keep up with growing demand. Current state of the art algorithms can handle some routine datasets but struggle when datasets grow above a certain size. We present a training deep learning via neural networks on non-linear dimensionality reduction, in particular t-distributed stochastic neighbour embedding (t-SNE), to overcome prior limitations of these methods. One Sentence SummaryAnalysis of prohibitively large datasets by combining deep learning via neural networks with non-linear dimensionality reduction.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.