Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow
Rosen, J. D.; Vasanthakumari, A. D.; Salomon, K.; de Lange, N.; Dash, P. M.; Keukeleire, P.; Hassan, A.; Barrera, A.; Kircher, M.; Love, M. I.; Schubach, M.
Show abstract
As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of non-coding genomic variation remains a major challenge due to the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively Parallel Reporter Assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, and systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing and visualization. Using diverse MPRA datasets, we characterize technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FinaleDB: a browser and database of cell-free DNA fragmentation patterns 95%
- Variant Library Annotation Tool (VaLiAnT): an oligonucleotide library design and annotation tool for Saturation Genome Editing and other Deep Mutational Scanning experiments 95%
- Flexiplex: A versatile demultiplexer and search tool for omics data 95%
Similar papers in this journal
- Improved characterization of single-cell RNA-seq libraries with paired-end avidity sequencing 96%
- Long-Read Structural and Epigenetic Profiling of a Kidney Tumor-Matched Sample with Nanopore Sequencing and Optical Genome Mapping 95%
- MUFFIN : A suite of tools for the analysis of functional sequencing data 94%
Similar papers in this journal
- txtools: an R package facilitating analysis of RNA modifications, structures, and interactions 95%
- PCLIPtools: A Robust Framework for Identifying RNA-Protein Interaction Sites from PAR-CLIP experiments. 95%
- Direct RNA sequencing (RNA004) allows for improved transcriptome assessment and near real-time tracking of methylation for medical applications 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.