A statistical model for quantitative analysis of single-molecule footprinting data
Ozonov, E. A.; Gaspa-Toneu, L.; Peters, A.
Show abstract
The binding of sequence-specific TFs (TF) to genomic DNA is fundamental to gene regulation. Emerging single-molecule footprinting (SMF) technologies such as the NOMe-seq and Fiber-seq assays offer unique opportunities for acquiring quantitative information about binding states of TFs and nucleosomes at single-DNA-molecule resolution. Contrasting bulk epigenomic profiling methodologies, SMF enables better molecular characterization of inherently stochastic processes of protein-DNA interactions. Despite the many advantages that SMF technologies bring for studying mechanisms of gene regulation, rigorous statistical models for the analysis of datasets generated using these technologies are still missing. Here, we introduce a novel statistical framework designed for inference of footprint lengths and predictions of footprint positions for unbiased quantitative analysis and interpretation of SMF datasets. We carried out comprehensive computational simulations of SMF experiments and identified experimental parameters that are critical to footprint detection. Finally, we demonstrate the power of this statistical approach for the analysis of genome-wide and amplicon-based NOMe-seq datasets generated for mouse embryonic stem cells.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deciphering the 3D genome organization across species from Hi-C data 97%
- Determination of human DNA replication origin position and efficiency reveals principles of initiation zone organisation 96%
- ChromPolymerDB: A High-Resolution Database of Single-Cell 3D Chromatin Structures for Functional Genomics 96%
Similar papers in this journal
- Chromatin information content landscapes inform transcription factor and DNA interactions 98%
- Fine-mapping of nuclear compartments using ultra-deep Hi-C shows that active promoter and enhancer elements localize in the active A compartment even when adjacent sequences do not 96%
- Electrostatic properties of disordered regions control transcription factor search and pioneer activity 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.