Back

SEAHORSE: A Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments

Quackenbush, A.; Kolluri, J.; Biju, R.; Nhong, S.; DeConti, D.; Quackenbush, J.; Saha, E.

2025-08-21 bioinformatics
10.1101/2025.08.15.670514 bioRxiv
Show abstract

Large-scale, open-access data sets such as the Genotype Tissue Expression Project (GTEx) and The Cancer Genome Atlas (TCGA) include multi-omic data on large numbers of samples along with extensive clinical and phenotypic information. These datasets provide a unique opportunity to discover correlations among clinical and genomic data features that can lead to testable hypotheses and new discoveries. SEAHORSE (http://seahorse.networkmedicine.org/) is a web-based database and search tool for exploratory data analysis in which we have pre-computed statistical associations between available data elements. An easy-to-use user interface allows users to explore significant associations using tabulated summary statistics, data visualizations, and functional enrichment analyses (using RNA-seq data) for identified sets of genes. We describe the motivation and construction of SEAHORSE and demonstrate its utility by documenting several surprising association patterns observed across multiple tissues in GTEx and multiple different cancer types in TCGA.

Matching journals

The top 13 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.