RAmpSim: A Thermodynamic Simulator for Hybridization Capturein Metagenomic Sequencing
Zhang, A.; Boucher, C.; Noyes, N.; Yu, Y. W.
Show abstract
Hybridization (bait) capture combined with long-read sequencing enables targeted profiling within complex metagenomes but introduces systematic biases from bait multiplicity, sequence composition, and species abundance that existing simulators ignore. We present RAmpSim, a fast simulator that models bait-target hybridization and fragment capture using a thermodynamic nearest-neighbor energy model and Boltzmann-weighted sampling of binding sites. Fragments are generated through multinomial sampling parameterized by bait concentration, binding energy, and genomic abundance before being passed to existing long-read simulators for modeling platform-specific errors. Implemented in Rust, RAmpSim reproduces empirical within-genome coverage and cross-species enrichment patterns observed in capture-based metagenomic datasets. Compared to uniform-coverage baselines, RAmpSims simulated coverage distributions are up to an order of magnitude closer to real data with respect to earth movers distance. Classification analysis reveals high recall in classifying high coverage regions between simulated and experimental distributions while outperforming a uniform baseline. Supporting accurate benchmarking and bait-set evaluation, RAmpSim provides an interpretable, efficient framework for simulating capture-based metagenomic sequencing. Code Availabilityhttps://github.com/az002/RAmpSim.git
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sketching and sampling approaches for fast and accurate long read classification 95%
- PIPETS: A statistically informed, gene-annotation agnostic analysis method to study bacterial termination using 3'-end sequencing. 94%
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.