Back

A replicable and modular benchmark for long-read transcript quantification methods

Zare Jousheghani, Z.; Singh, N. P.; Patro, R.

2024-07-31 bioinformatics
10.1101/2024.07.30.605821 bioRxiv
Show abstract

We provide a replicable benchmark for long-read transcript quantification, and evaluate the performance of some recently-introduced tools on several synthetic long-read RNA-seq datasets. This benchmark is designed to allow the results to be easily replicated by other researchers, and the structure of the underlying Snakemake workflow is modular to make the addition of new tools or new data sets relatively easy. In analyzing previously assessed simulations, we find discrepancies with recently-published results. We also demonstrate that the robustness of certain approaches hinge critically on the quality and "cleanness" of the simulated data. AvailabilityThe Snakemake scripts for the benchmark are available at https://github.com/COMBINE-lab/lr_quant_benchmarks, the data used as input for the benchmarks (reference sequences, annotations, and simulated reads) are available at https://doi.org/10.5281/zenodo.13130623.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.