Analysis and Augmentation of Small Datasets with Unsupervised Machine Learning
Dolgikh, S.
Show abstract
Analysis of small datasets presents a number of essential challenges not in the least due to insufficient sampling of characteristic patterns in the data making confident conclusions about the unknown distribution elusive and resulting in lower statistical confidence and higher error. In this work, a novel approach to augmentation of small datasets is proposed based on an ensemble of neural network models of unsupervised generative self-learning. Applying generative learning with an ensemble of individual models allowed to identify stable clusters of data points in the latent representations of the observable data. Several techniques of augmentation based on identified latent cluster structure were applied to produce new data points and enhance the dataset. The proposed method can be used with small and extremely small datasets to identify characteristics patterns, augment data and in some cases, improve accuracy of classification in the scenarios with strong deficit of labels.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Decoding Clinical Biomarker Space of COVID-19: Exploring Matrix Factorization-based Feature Selection Methods 96%
- A machine-learning Approach for Stress Detection Using Wearable Sensors in Free-living Environments 96%
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 94%
Similar papers in this journal
- A Robust Spike Sorting Method based on the Joint Optimization of Linear Discrimination Analysis and Density Peaks 94%
- A new, simple method of describing COVID-19 trajectory and dynamics in any country based on Johnson Cumulative Distribution Function fitting 94%
- A novel interpretable deep transfer learning combining diverse learnable parameters for improved T2D prediction based on single-cell gene regulatory networks 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.