Beyond Blacklists: A Critical Assessment of Exclusion Set Generation Strategies and Alternative Approaches
Wall, B. P. G.; Ogata, J. D.; Nguyen, M.; McClay, J. L.; Harrell, C.; Dozmorov, M. G.
Show abstract
Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Tools like the widely used Blacklist software facilitate this process, but their algorithmic details and parameter choices are not always clearly documented, affecting reproducibility and biological relevance. We examined the Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to variability in input data, aligner choice, and read length. We also identified and addressed a coding issue that led to over-annotation of high-signal regions. We further explored the use of "sponge" sequences--unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA--as an alternative approach. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets and suggest that sponge sequences offer a flexible, alignment-guided strategy for reducing artifacts and improving functional genomics analyses.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Personalized and graph genomes reveal missing signal in epigenomic data 97%
- RADAR: Differential analysis of MeRIP-seq data with a random effect model 95%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 95%
Similar papers in this journal
- Assessment of human diploid genome assembly with 10x Linked-Reads data 97%
- RepeatFiller newly identifies megabases of aligning repetitive sequences and improves annotations of conserved non-exonic elements 95%
- LRTK: A platform agnostic toolkit for linked-read analysis of both human genomes and metagenomes 94%
Similar papers in this journal
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 96%
- Allo: Accurate allocation of multi-mapped reads enables regulatory element analysis at repeats 96%
- Ultra-low input single tube linked-read library method enables short-read NGS systems to generate highly accurate and economical long-range sequencing information for de novo genome assembly and haplotype phasing 95%
Similar papers in this journal
- Identification and Utilization of Copy Number Information for Correcting Hi-C Contact Map of Cancer Cell Line 96%
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 96%
- SpectralTAD: an R package for defining a hierarchy of Topologically Associated Domains using spectral clustering 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.