Effective design and inference for cell sorting and sequencing based massively parallel reporter assays
Gilliot, P.-A.; Gorochowski, T. E.
Show abstract
The ability to measure the phenotype of millions of different genetic designs using Massively Parallel Reporter Assays (MPRAs) has revolutionised our understanding of genotype-to-phenotype relationships and opened avenues for data-centric approaches to biological design. However, our knowledge of how best to design these costly experiments and the effect that our choices have on the quality of the data produced is lacking. Here, we tackle this issue by developing FORE-CAST, a Python package that supports the accurate simulation of cell-sorting and sequencing based MPRAs and robust maximum like-lihood based inference of genetic design function from MPRA data. We use FORECASTs capabilities to reveal rules for MPRA experimental design that help ensure accurate genotype-to-phenotype links and show how the simulation of MPRA experiments can help us better understand the limits of prediction accuracy when this data is used for training deep learning based classifiers. As the scale and scope of MPRAs grows, tools like FORECAST will help ensure we make informed decisions during their development and the most of the data produced.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- On the discovery of population-specific state transitions from multi-sample multi-condition single-cell RNA sequencing data 95%
- Single-cell DNA replication dynamics in genomically unstable cancers 95%
- Statistical modeling, estimation, and remediation of sample index hopping in multiplexed droplet-based single-cell RNA-seq data 95%
Similar papers in this journal
Similar papers in this journal
- Predicting Composition of Genetic Circuits with Resource Competition: Demand and Sensitivity 93%
- Meeting Measurement Precision Requirements for Effective Engineering of Genetic Regulatory Networks 93%
- The Synthesis Success Calculator: Predicting the Rapid Synthesis of DNA Fragments with Machine Learning 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.