excluderanges: exclusion sets for T2T-CHM13, GRCm39, and other genome assemblies
Ogata, J. D.; Mu, W.; Davis, E. S.; Xue, B.; Harrell, J. C.; Sheffield, N. C.; Phanstiel, D. H.; Love, M. I.; Dozmorov, M. G.
Show abstract
SummaryExclusion regions are sections of reference genomes with abnormal pileups of short sequencing reads. Removing reads overlapping them improves biological signal, and these benefits are most pronounced in differential analysis settings. Several labs created exclusion region sets, available primarily through ENCODE and Github. However, the variety of exclusion sets creates uncertainty which sets to use. Furthermore, gap regions (e.g., centromeres, telomeres, short arms) create additional considerations in generating exclusion sets. We generated exclusion sets for the latest human T2T-CHM13 and mouse GRCm39 genomes and systematically assembled and annotated these and other sets in the excluderanges R/Bioconductor data package, also accessible via the BEDbase.org API. The package provides unified access to 82 GenomicRanges objects covering six organisms, multiple genome assemblies and types of exclusion regions. For human hg38 genome assembly, we recommend hg38.Kundaje.GRCh38_unified_blacklist as the most well-curated and annotated, and sets generated by the Blacklist tool for other organisms. Availability and implementationhttps://bioconductor.org/packages/excluderanges/ ContactMikhail G. Dozmorov (mdozmorov@vcu.edu) Supplementary informationPackage website: https://dozmorovlab.github.io/excluderanges/
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Pushing the limits of HiFi assemblies reveals centromere diversity between two Arabidopsis thaliana genomes 95%
- Quality-controlled R-loop meta-analysis reveals the characteristics of R-Loop consensus regions 94%
- PCLIPtools: A Robust Framework for Identifying RNA-Protein Interaction Sites from PAR-CLIP experiments. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.