A replicable and modular benchmark for long-read transcript quantification methods
Zare Jousheghani, Z.; Singh, N. P.; Patro, R.
Show abstract
We provide a replicable benchmark for long-read transcript quantification, and evaluate the performance of some recently-introduced tools on several synthetic long-read RNA-seq datasets. This benchmark is designed to allow the results to be easily replicated by other researchers, and the structure of the underlying Snakemake workflow is modular to make the addition of new tools or new data sets relatively easy. In analyzing previously assessed simulations, we find discrepancies with recently-published results. We also demonstrate that the robustness of certain approaches hinge critically on the quality and "cleanness" of the simulated data. AvailabilityThe Snakemake scripts for the benchmark are available at https://github.com/COMBINE-lab/lr_quant_benchmarks, the data used as input for the benchmarks (reference sequences, annotations, and simulated reads) are available at https://doi.org/10.5281/zenodo.13130623.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Tailored machine learning models for functional RNA detection in genome-wide screens 95%
- DiffSegR: An RNA-Seq data driven method for differential expression analysis using changepoint detection 94%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Alignment and mapping methodology influence transcript abundance estimation 96%
- Enhancing transcriptome expression quantification through accurate assignment of long RNA sequencing reads with TranSigner 95%
- Identifying and quantifying isoforms from accurate full-length transcriptome sequencing reads with Mandalorion 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.