Back

SummArIzeR: Simplifying cross-database enrichment result clustering and annotation via large language models

Brinkmann, M.; Bonelli, M.; Tosevska, A.

2025-06-01 bioinformatics
10.1101/2025.05.28.656331 bioRxiv
Show abstract

MotivationEnrichment analysis across multiple databases often results in a high level of redundancy due to overlapping terms, complicating the interpretation of biological data. To address this, we developed SummArIzeR, an R package to cluster and annotate enrichment results across multiple databases, enabling fast, intuitive interpretation and comparison across multiple conditions. SummArIzeR clusters enrichment results based on shared genes, calculates a pooled p-value for each cluster and facilitates the cluster annotation using large-language models. It further allows an easyly interpretable vizualisisation of the results. ResultsCompared to existing tools, SummArIzeR provides unbiased and fast cluster annotation using large language models. We demonstrate that SummArIzeR achieves clustering comparable to manual curation while offering superior grouping based on shared underlying genes. Availability and ImplementationThe SummArIzeR package is available as an open-source R package, with a comprehensive user manual provided in its GitHub repository: https://github.com/bonellilab/SummArIzeR.

Published in Bioinformatics (predicted rank #4) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.