Solving Anscombe's Quartet using a Transfer Learning Approach
Bu, K.; Clemente, J. C.
Show abstract
Analysis of high-dimensional datasets often involves usage of summary statistics, one of which is the correlation coefficient. These values are then used to inform downstream analysis, whether in feature selection or in subsequent construction of networks and heatmaps. Condensing pairwise scatterplots into these singular values however, often results in a loss of information. Originally proposed by F. J. Anscombe in his famous Anscombes Quartet, this phenomenon has been canonically used to demonstrate the importance of plotting and the limitations of summary statistics such as correlation or variance [F.J. Anscombe, (1973) American Statistician. 27 (1), 17-21]. While numerous methods exist for the generation of visually distinct datasets that share similar summary statistics, the converse has not been extensively studied. To address this gap, we propose ICLUST (Image CLUSTering), an image classifier tool that can visually distinguish correlations with similar summary statistics in simulations and identify meaningful clusters in real data. Such a tool can potentially benefit those performing exploratory analysis or feature selection in a complementary fashion by identifying relationships between variables that traditional summary metrics cannot provide. Significance StatementDistilling large-scale, multidimensional datasets via analysis of pairwise relationships often employs a single value to describe the relationship between variables. However, as demonstrated through simulations, such summarization fails to retain the nuances of the data. Characteristics such as the type of relationship (linear versus nonlinear, etc.) and the spread of the data are commonly lost when using correlations. Here we propose a transfer learning framework, borrowing from image clustering and classification software, to visually classify graphs. We apply our method towards separation of scatterplots with similar correlation statistics but visually distinctive patterns in both simulations and real data, demonstrating its broad applicability.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 95%
- Haisu: Hierarchical Supervised Nonlinear Dimensionality Reduction 94%
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 94%
Similar papers in this journal
Similar papers in this journal
- An Assistive Computer Vision Tool to Automatically Detect Changes in Fish Behavior In Response to Ambient Odor 94%
- Novel AI-powered computational method using tensor decomposition for identification of common optimal bin sizes when integrating multiple Hi-C datasets 94%
- Visibility Graph Based Community Detection for Biological Time Series 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.