SummArIzeR: Simplifying cross-database enrichment result clustering and annotation via large language models
Brinkmann, M.; Bonelli, M.; Tosevska, A.
Show abstract
MotivationEnrichment analysis across multiple databases often results in a high level of redundancy due to overlapping terms, complicating the interpretation of biological data. To address this, we developed SummArIzeR, an R package to cluster and annotate enrichment results across multiple databases, enabling fast, intuitive interpretation and comparison across multiple conditions. SummArIzeR clusters enrichment results based on shared genes, calculates a pooled p-value for each cluster and facilitates the cluster annotation using large-language models. It further allows an easyly interpretable vizualisisation of the results. ResultsCompared to existing tools, SummArIzeR provides unbiased and fast cluster annotation using large language models. We demonstrate that SummArIzeR achieves clustering comparable to manual curation while offering superior grouping based on shared underlying genes. Availability and ImplementationThe SummArIzeR package is available as an open-source R package, with a comprehensive user manual provided in its GitHub repository: https://github.com/bonellilab/SummArIzeR.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cluefish: mining the dark matter of transcriptional data series with over-representation analysis enhanced by aggregated biological prior knowledge 95%
- DecoPath: A web application for decoding pathway enrichment analysis 94%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 94%
Similar papers in this journal
Similar papers in this journal
- HUMESS: Integrating Quantitative Transcriptomic Analysis and Metabolic Modeling to Unveil Condition-Specific Gene Signatures 94%
- Sub-Cluster Identification through Semi-SupervisedOptimization of Rare-cell Silhouettes (SCISSORS) in Single-Cell Sequencing 93%
- Pathway Volcano: An interactive tool for pathway guided visualization of differential expression data 93%
Similar papers in this journal
- iTraNet: A Web-Based Platform for integrated Trans-Omics Network Visualization and Analysis 94%
- SPECK: An Unsupervised Learning Approach for Cell Surface Receptor Abundance Estimation for Single Cell RNA-Sequencing Data 93%
- scExplorer: A Comprehensive Web Server for Single-Cell RNA Sequencing Data Analysis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.