Back

Mango: Unearthing Patterns in Large-Scale Biological Data Through Interactive Correlation Analysis

Branders, S.; Grabherr, M. G.; Ahmad, R.

2025-10-07 bioinformatics
10.1101/2025.10.06.680608 bioRxiv
Show abstract

Integrating different types of biological data is often challenging due to the presence of both numerical and categorical data. This complexity makes it harder to evaluate causal biological effects, especially when confounders like population structure, sampling methods, or multi-omics integration can lead to incorrect conclusions. We introduce Mango, an interactive correlation browser designed for visually exploring any tabular data type, using a novel algorithm to correlate numerical and categorical data, regardless of their distribution, called Median-Ranked Label Encoding. Our results on genomic and transcriptomic datasets demonstrate that these correlations can effectively distinguish between biases and causal relationships in large-scale data.

Published in BMC Genomics (predicted rank #5) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.