Mango: Unearthing Patterns in Large-Scale Biological Data Through Interactive Correlation Analysis
Branders, S.; Grabherr, M. G.; Ahmad, R.
Show abstract
Integrating different types of biological data is often challenging due to the presence of both numerical and categorical data. This complexity makes it harder to evaluate causal biological effects, especially when confounders like population structure, sampling methods, or multi-omics integration can lead to incorrect conclusions. We introduce Mango, an interactive correlation browser designed for visually exploring any tabular data type, using a novel algorithm to correlate numerical and categorical data, regardless of their distribution, called Median-Ranked Label Encoding. Our results on genomic and transcriptomic datasets demonstrate that these correlations can effectively distinguish between biases and causal relationships in large-scale data.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GeneSetCluster 2.0: a comprehensive toolset for summarizing and integrating gene-sets analysis 95%
- ChiMera: An easy to use pipeline for Bacterial Genome Based Metabolic Network Reconstruction, Evaluation and Visualization 94%
- Hypercluster: a flexible tool for parallelized unsupervised clustering optimization 93%
Similar papers in this journal
- VIBES: A Workflow for Annotating and Visualizing Viral Sequences Integrated into Bacterial Genomes 94%
- ResistoXplorer: a web-based tool for visual, statistical and exploratory data analysis of resistome data 94%
- Dashboard-style interactive plots for RNA-seq analysis are R Markdown ready with Glimma 2.0 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.