Meta-analysis of RNA-seq gene fusion detection tools: performance and variability across benchmarks
Ostrowska, I.; Gambin, T.
Show abstract
In this study, we conducted a comprehensive meta-analysis of ten publicly available benchmark tests evaluating the performance of gene fusion detection tools using RNA-seq data. Our analysis focused on key performance metrics, including sensitivity, precision and F1 scores. We evaluated how the tools perform in different data sets. We examined the impact of data set characteristics such as sample type (real or simulated) and read length, as well as additional sample and sequencing parameters on the results. In addition to evaluating performance, we analyzed the organization and design of the benchmark tests, highlighting good practices such as clear descriptions of the datasets, detailed instrument parameters and transparent methodologies. However, we also identified common pitfalls, including insufficient reproducibility information, limited diversity of datasets, and the lack of widely accepted gold standard datasets. These limitations make it difficult to consistently evaluate tools and compare across benchmarks. By synthesizing these findings, we make recommendations for future benchmark projects, emphasizing the need for standardization, increased transparency and the development of robust truth sets. This study aims to help the community create more reliable and reproducible benchmark tests, ultimately accelerating the development and evaluation of gene fusion detection tools for clinical and research applications.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 93%
- labelSeg: segment annotation for tumor copy number alteration profiles 92%
- SAMURAI: Shallow Analysis of copy nuMber alterations Using a Reproducible And Integrated bioinformatics pipeline 92%
Similar papers in this journal
- Differential Expression Analysis with InMoose, the Integrated Multi-Omic Open-Source Environment in Python 93%
- Performance analysis of conventional and AI-based variant callers using short and long reads 93%
- Probabilistic modeling methods for cell-free DNA methylation based cancer classification 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- eDAVE - extension of GDC Data Analysis, Visualization, and Exploration Tools 92%
- UMI-Gen: a UMI-based reads simulator for variant calling evaluation in paired-end sequencing NGS libraries 92%
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.