Nature Biotechnology
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match Nature Biotechnology's content profile, based on 172 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Walton, R. T.; Qin, Y.; Blainey, P. C.
Show abstract
Forward genetic screens seek to dissect complex biological systems by systematically perturbing genetic elements and observing the resulting phenotypes. While standard screening methodologies introduce individual perturbations, multiplexing perturbations improves the performance of single-target screens and enables combinatorial screens for the study of genetic interactions. Current tools for multiplexing perturbations are limited by technical challenges and do not offer compatibility across diverse screening methodologies, including enrichment, single-cell sequencing, and optical pooled screens. Here, we report the development of CROPseq-multi (CSM), a CROPseq1-inspired lentiviral system to multiplex Streptococcus pyogenes (Sp) Cas9-based perturbations with versatile readout compatibility and high performance for both perturbation and barcode identification. CSM has equivalent per-guide activity to CROPseq and low lentiviral recombination frequencies. Dual-guide CSM libraries are constructed in a single, facile molecular cloning step that facilitates the use of unique molecular identifiers. CSM is compatible with enrichment screening methodologies, single-cell RNA-sequencing readouts, and optical pooled screens. For optical pooled screens, an optimized and multiplexed in situ detection protocol improves barcode counts 10-fold (for mRNA detection), enables detection of recombination events, and reduces the number of sequencing cycles required for decoding by 3-fold relative to CROPseq. CROPseq-multi-v2 (CSMv2) adds compatibility for detection methods based on T7 RNA polymerase in vitro transcription2-5. CSM provides a single system for CRISPR screens that is compatible with individual and combinatorial perturbations, diverse SpCas9-based perturbation technologies, and multiple high-content, single-cell phenotypic readouts.
Cheng, A. P.; Rusinek, I.; Sossin, A.; Widman, A. J.; Meiri, E.; Krieger, G.; Hirschberg, O.; Tov, D. S.; Gilad, S.; Jaimovich, A.; Barad, O.; Avaylon, S.; Rajagopalan, S.; Potenski, C.; Prieto, T.; Yuan, D. J.; Furatero, R.; Runnels, A.; Costa, B. M.; Shoag, J. E.; Al Assaad, M.; Sigouros, M.; Manohar, J.; King, A.; Wilkes, D.; Otilano, J.; Malbari, M. S.; Elemento, O.; Mosquera, J. M.; Altorki, N. K.; Saxena, A.; Callahan, M. K.; Robine, N.; Germer, S.; Evrony, G.; Faltas, B. M.; Landau, D. A.
Show abstract
Distinguishing real biological variation in the form of single-nucleotide variants (SNVs) from errors is a major challenge for genome sequencing technologies. This is particularly true in settings where SNVs are at low frequency such as cancer detection through liquid biopsy, or human somatic mosaicism. State-of-the-art molecular denoising approaches for DNA sequencing rely on duplex sequencing, where both strands of a single DNA molecule are sequenced to discern true variants from errors arising from single stranded DNA damage. However, such duplex approaches typically require massive over-sequencing to overcome low capture rates of duplex molecules. To address these challenges, we introduce paired plus-minus sequencing (ppmSeq) technology, in which both DNA strands are partitioned and clonally amplified on sequencing beads through emulsion PCR. In this reaction, both strands of a double-stranded DNA molecule contribute to a single sequencing read, allowing for a duplex yield that scales linearly with sequencing coverage across a wide range of inputs (1.8-98 ng). We benchmarked ppmSeq against current duplex sequencing technologies, demonstrating superior duplex recovery with ppmSeq, with a rate of 44%{+/-}5.5% (compared to [~]5-11% for leading duplex technologies). Using both genomic as well as cell-free DNA, we established error rates for ppmSeq, which had residual SNV detection error rates as low as 7.98x10-8 for gDNA (using an end-repair protocol with dideoxy nucleotides) and 3.5x10-7{+/-}7.5x10-8 for cell-free DNA. To test the capabilities of ppmSeq for error-corrected whole-genome sequencing (WGS) for clinical application, we assessed circulating tumor DNA (ctDNA) detection for disease monitoring in cancer patients. We demonstrated that ppmSeq enables powerful tumor-informed ctDNA detection at concentrations of 10-4 across most cancers, parts per million sensitivity in cancers with high mutation burden, and further increased sensitivity with higher sequencing depth. We then leveraged genome-wide trinucleotide mutation patterns characteristic of urothelial (APOBEC3-related and platinum exposure-related signatures) and lung (tobacco-exposure-related signatures) cancers to perform tumor-naive ctDNA detection, showing that ppmSeq can identify a disease-specific signal in plasma cell-free DNA without a matched tumor, and that this signal correlates with imaging-based disease metrics. Altogether, ppmSeq provides an error-corrected, cost-efficient and scalable approach for high-fidelity WGS that can be harnessed for challenging clinical applications and emerging frontiers in human somatic genetics where high accuracy is required for mutation identification.
Nilsson, A.; Sporre, E.; Schulte, D.; Snijder, J.; Edfors, F.; Käll, L.
Show abstract
Reading a proteins sequence from tandem mass spectra without a reference is limited by single-spectrum accuracy, most acutely across the hypervariable complementaritydetermining regions of antibodies. Broadly specific proteases tile a protein with long, overlapping peptides, so every residue is covered by many independent de novo reads. borgonovo assembles their per-step probability profiles into a reference-free per-residue consensus, seeding templates from mass-closure-consistent reads and recruiting the rest by substitution- tolerant alignment and per-column voting. Re-decoding each spectrum with a prior from its consensus position lifts amino acid accuracy on placed spectra from 0.80 to 0.87. On the therapeutic antibody trastuzumab, nine proteases cover its heavy and light chains completely at 0.88 fixed-window identity, and 0.93 on the pruned assembly once local indels are accommodated. Applied unchanged to five secretome proteins and trastuzumab with three proteases, it reaches 0.87 mean fixed-window identity over 82% coverage. borgonovo is open source and works with most de novo sequencers, so redundant digestion turns any of them into a protein sequencer where no reference exists.
Piscopio, R. A.; Chialastri, A.; Wang, C.; Godzik, M.; Heom, K. A.; Wang, W.; Wilson, M. Z.; Dey, S. S.
Show abstract
The organization of cells within a tissue plays a critical role in tuning cellular function. Several methods have recently been developed to capture the transcriptome of cells while retaining spatial information. However, these genome-wide sequencing methods typically lack the spatial resolution of individual cells and are confined to quantifying positional information within predefined lattice locations, thereby failing to capture large sections of a tissue outside these regions. Further, these methods are generally limited to profiling fixed cells with reduced mRNA capture efficiency compared to standard scRNA-seq. In addition, existing methods lack modularity and cross-platform compatibility, thereby limiting most of these techniques from jointly profiling the epigenetic and transcriptomic state of individual cells. To overcome these limitations, we present scSTAMP-seq (single-cell Spatial Transcriptomic And Multiomic Profiling), an approach that employs cholesterol-tagged photolabile oligonucleotides that incorporate into cell membranes, enabling us to "stamp" the position of cells using spatially imposed light gradients prior to tissue dissociation and single-cell sequencing. Applied to live cells, scSTAMP-seq efficiently captures spatially resolved single-cell transcriptomes at high resolution for all cells within a field of view. Further, we demonstrate that light patterning enables dynamic spatial resolution, including the ability to map the position of individual cells. Finally, we show that scSTAMP-seq is modular and can be seamlessly integrated with various downstream single-cell sequencing technologies. We demonstrate this by performing scRNA-seq using plate- and droplet-based methods, and by performing joint epigenome and transcriptome sequencing from the same cell while preserving positional information. Collectively, these results demonstrate that scSTAMP-seq is a sensitive and high-throughput technology for mapping single-cell transcriptomes and epigenomes at the spatial resolution of individual cells.
Lalanne, J.-B.; Mich, J. K.; Huynh, C.; Hunker, A.; McDiarmid, T. A.; Levi, B. P.; Ting, J. T.; Shendure, J.
Show abstract
Adeno-associated viruses (AAVs) have emerged as the foremost gene therapy delivery vehicles due to their versatility, durability, and safety profile. Here we demonstrate extensive chimerism, manifesting as pervasive barcode swapping, among complex AAV libraries that are packaged as a pool. The observed chimerism is length- and homology-dependent but capsid-independent, in some cases affecting the majority of packaged AAV genomes. These results have implications for the design and deployment of functional AAV libraries in both research and clinical settings.
Honigfort, D.; Belda-Ferre, P.; White, D.; Sundararajan, K.; Dawood, M.; LeVieux, J.; Moreno, J.; Qi, X.; Metcalfe, K.; Naranbat, D.; Altomare, A.; Thompson, C.; Perez, C.; Lajoie, B.; Kwon, H.; Bhadha, P.; Rammel, T.; Rabalais, J.; Kellinger, M.; Kruglyak, S.; Arslan, S.; Previte, M.
Show abstract
We present a platform that directly sequences single guide RNAs and endogenous 3'UTRs in fixed cells while simultaneously measuring protein abundance and cellular morphology. We demonstrate platform capability by performing optical pooled screening of CRISPR-perturbed lung cancer cells. This approach unites direct in-sample RNA sequencing with complementary phenotypic readouts, enabling comprehensive, scalable, and functional genomics analyses within a single experiment.
Ghaddar, B.; Blaser, M. J.; De, S.
Show abstract
We developed SAHMI, a computational resource to identify truly present microbial nucleic acids and filter contaminants and spurious false-positive taxonomic assignments from standard transcriptomic sequencing of mammalian tissues. In benchmark studies, SAHMI correctly identifies known microbial infections present in diverse tissues. The application of SAHMI to single-cell and spatial genomic data enables co-detection of somatic cells and microorganisms and joint analysis of host-microbiome ecosystems.
Langley, J.; Baudrier, L.; Curry, J.; Narta, K.; Todesco, H. M.; Potts, K.; Morrissy, S.; Mahoney, D. J.; Billon, P.
Show abstract
Engineered virus-like particles (eVLPs) enable transgene-free ribonucleoprotein delivery for genome editing applications, yet optimized delivery strategies for high-throughput applications remain unexplored. Prime editing enables precise genomic modifications but suffers from limited efficiency that constrains its widespread adoption. Here, we present PRIME-VLP (Progressive Repeated Infections for Maximized Editing via Virus-Like Particles), a delivery strategy that enhances prime editing efficiency for both targeted genome engineering and high-throughput prime editing screening. PRIME-VLP leverages the temporal dynamics of eVLP-mediated editing through multiple sequential transductions with sub-saturating eVLP doses delivered at optimal intervals. This approach achieves 1.5 to 2.8-fold improvements in editing efficiency across diverse genomic targets and cell types. PRIME-VLP maintains high specificity without increasing off-target effects, compromising cellular viability or causing transcriptional perturbations. By decoupling pegRNA and editor delivery through pegRNA-free eVLPs, PRIME-VLP enables pooled prime editing screens, circumventing transgene silencing limitations of conventional lentiviral-based screens. Using a 6,000-pegRNA library targeting TP53, PRIME-VLP achieved 2.8-fold higher editing efficiency and improved reproducibility compared to conventional lentiviral delivery. An eVLP-based screen identified functional TP53 loss-of-function variants that confer resistance to MDM2 inhibition by Nutlin-3. This work expands the versatility of eVLPs beyond their current in vivo therapeutic applications, demonstrating their promise for high-throughput functional genomics by overcoming the delivery limitations of lentiviral systems.
Almogy, G.; Pratt, M.; Oberstrass, F.; Lee, L.; Mazur, D.; Beckett, N.; Barad, O.; Soifer, I.; Perelman, E.; Etzioni, Y.; Sosa, M.; Jung, A.; Clark, T.; Trepagnier, E.; Lithwick-Yanai, G.; Pollock, S.; Hornung, G.; Levy, M.; Coole, M.; Howd, T.; Shand, M.; Farjoun, Y.; Emery, J.; Hall, G.; Lee, S. K.; Sato, T.; Magner, R.; Low, S.; Bernier, A.; Gandi, B.; Stohlman, J.; Nolet, C.; Donovan, S.; Blumenstiel, B.; Cipicchio, M.; Dodge, S.; Banks, E.; Lennon, N.; Gabriel, S.; Lipson, D.
Show abstract
We introduce a massively parallel novel sequencing platform that combines an open flow cell design on a circular wafer with a large surface area and mostly natural nucleotides that allow optical end-point detection without reversible terminators. This platform enables sequencing billions of reads with longer read length ([~]300bp) and fast runs times (<20hrs) with high base accuracy (Q30 > 85%), at a low cost of $1/Gb. We establish system performance by whole-genome sequencing of the Genome-In-A-Bottle reference samples HG001-7, demonstrating high accuracy for SNPs (99.6%) and Indels in homopolymers up to length 10 (96.4%) across the vast majority (>98%) of the defined high-confidence regions of these samples. We demonstrate scalability of the whole-genome sequencing workflow by sequencing an additional 224 selected samples from the 1000 Genomes project achieving high concordance with reference data.
Khafizov, R.; Piazza, E.; Cui, Y.; Patrick, M.; Metzger, E.; McGuire, D.; Dunaway, D.; Danaher, P.; Hoang, M. L.; Grootsky, A.; Vandenberg, M.; He, S.; Liu, R.; McKean, M.; Rhodes, M.; Beechem, J. M.
Show abstract
Single-cell RNA-seq revolutionized single-cell biology, by providing a complete whole transcriptome view of individual cells. Regrettably, this was accomplished only for individual, tissue-dissociated cells. High-plex spatial biology has begun to recover the x, y, and z-coordinates of single-cells, but typically at the expense of far less than whole transcriptome coverage. To solve this problem, Bruker Spatial Biology has accomplished a commercial-grade panel (CosMx(R) Spatial Molecular Imager Whole Transcriptome Panel (WTx)), using 37,872 imaging barcodes, capable of sub-cellular imaging of the entire human protein-coding transcriptome. The imaging barcodes are encoded with 156 bits of information (4 on-cycles and 35 dark-cycles per code), at a Hamming Distance of 4 from each other to achieve a very low false-code detection. Key to achieving this high-plex capability was the ability to manufacture imaging barcodes that require no in-tissue amplification (every barcode is manufactured under GMP to contain exactly 30 fluorescent dyes) and uniform, size-exclusion purified, extremely small imaging barcodes ([~] 20 nm). A detailed study of six different human FFPE tissue types was performed (Colon, Pancreas, Hippocampus, Skin, Breast, Kidney), yielding over 5.4 billion transcripts from 2.7 million cells. We counted over 1,550 transcripts-per-cell on average and observed 900 unique genes per cell (measured as the median). Single fixed-cells containing well over 10,000 subcellularly imaged transcripts were accomplished. Advancing single-cell imaging to the whole transcriptome level opens a single unified approach to accomplish essentially all single-cell experiments, both imaging and non-imaging. Depending upon the sample type (e.g. fixed-cells, organoids, tissue sections, etc.), the transcripts per cell and genes per cell measured using the whole transcriptome panel often exceeds that obtained by the highest-resolution single-cell RNA-seq, can be performed on a single 5 {micro}m FFPE tissue section, with no dissociation bias (every cell is counted). Pathway analysis within the tumor bed of a colon adenocarcinoma sample found evidence of enrichment in pathways suggestive of an aggressive tumor type, and localized ligand-receptor analysis showed spatially restricted patterns related to adhesion, migration, and proliferation. The high-dimensional whole transcriptome data is streamed directly to a cloud-based Spatial Informatics Platform, allowing for the scalable processing of millions-of-single-cells and billions-of-transcripts per operation. The WTx data are combined with high-resolution antibody-based cell-morphology imaging and data-driven machine-learning cell segmentation algorithms, to generate the most complete view of single cell and sub-cellular spatial biology that has ever been obtained.
Loyfer, N.; Magenheim, J.; Darwish, A.; Isaac, S.; Ganbat, J.; Babikir, H.; Jhutty, A.; Wan, J.; Bayes-Genis, A.; Revuelta-Lopez, E.; Eden, A.; Solanki, R.; Dor, Y.; Kaplan, T.
Show abstract
Enzymatic methylation sequencing (EM-seq) converts unmethylated cytosines to uracils while preserving DNA integrity, making it attractive for liquid biopsies. Here we report a reproducible fragment-level over-conversion error in EM-seq, which is not observed in bisulfite-based conversion or in Oxford Nanopore sequencing. While bisulfite and nanopore errors occur at sporadic CpGs, EM-seq generates molecules that appear fully unmethylated, introducing a false-positive background signal that severely limits deconvolution specificity in cfDNA analysis.
Bradu, A.; Blair, J. D.; Grabski, I. N.; Mascio, I.; Lee, J.; McCormick, C.; Satija, R.
Show abstract
CRISPR-based screening combined with single-cell sequencing (i.e. Perturb-seq) enables systematic mapping of genetic perturbations to molecular phenotypes. While Perturb-seq is well-suited to profile targeted subsets of regulators, scaling to genome-wide screens presents substantial cost and throughput challenges. Here we introduce VIPerturb-seq, a platform to facilitate routine genome-wide Perturb-seq experiments using probe-based detection workflows. We describe a split probe strategy for detection of genome-wide CRISPR libraries in fixed cells that enables (i) optional support for phenotypic enrichment of Very Important Perturbations (VIP) prior to single-cell profiling, and (ii) compatibility with combinatorial indexing workflows to further improve Perturb-seq throughput by 50-fold. Using a genome-wide CRISPRi library (GuEST-List), we demonstrate VIPerturb-seq on two genome-wide screens representing both unbiased and phenotypically enriched workflows. Our results demonstrate how the sensitivity, scalability, and efficiency of VIPerturb-seq can enable both individual labs with targeted research questions and large data generation platforms aiming to construct virtual cells.
Goode, Z.; Tiedemann, E.; Ben Ameur, L.; Pavan, K.; Young, K.; Sek, M.; Nevue, A.; Zhu, J.; Houghton, J.; Fu, Y.; Boisvert, H.; Saunders, A.
Show abstract
Probe-based genomics technologies are extending molecular analysis into intact tissues and fixed cells, yet strategies to decode complex experimental conditions encoded in cellular RNA remain limited. Here we present a modular framework that integrates custom software tools with purpose-built cloning reagents to design, assemble, validate, and deploy combinatorial DNA barcodes. Combinatorial barcodes comprise spatially adjacent collections of known sequences, enabling millions of unique molecules to be efficiently distinguished using a limited set of probes. Our software tools integrate with optimized assembly plasmids and whole plasmid long-read sequencing for high-fidelity construction and structural validation of diverse combinatorial barcode architectures. Assembled barcode libraries are flexibly transferred into user-modified expression vectors to support diverse downstream experimental applications. We showcase the versatility of this framework by assembling two structurally distinct combinatorial barcode libraries, each containing millions of unique sequences. Following rabies virus-based delivery to the mouse brain, we validate in vivo decoding of a combinatorial barcode architecture capable of distinguishing ~16.3 million expressed RNAs through probe-based in situ sequencing. Our framework for flexible and accurate combinatorial barcode construction fills a technically demanding niche delivering cost-effective molecular reagents for multiplexed experimentation on current and evolving probe-based genomics platforms.
Kim, J.; Muller, R. Y.; Bondra, E. R.; Ingolia, N.
Show abstract
Genome-wide CRISPR screens have emerged as powerful tools for uncovering the genetic underpinnings of diverse biological processes. Incisive screens often depend on directly measuring molecular phenotypes, such as regulated gene expression changes, provoked by CRISPR-mediated genetic perturbations. Here, we provide quantitative measurements of transcriptional responses in human cells across genome-scale perturbation libraries by coupling CRISPR interference (CRISPRi) with barcoded expression reporter sequencing (CiBER-seq). To enable CiBER-seq in mammalian cells, we optimize the integration of highly complex, barcoded sgRNA libraries into a defined genomic context. CiBER-seq profiling of a nuclear factor kappa B (NF-{kappa}B) reporter delineates the canonical signaling cascade linking the transmembrane TNF-alpha receptor to inflammatory gene activation and highlights cell-type-specific factors in this response. Importantly, CiBER-seq relies solely on bulk RNA sequencing to capture the regulatory circuit driving this rapid transcriptional response. Our work demonstrates the accuracy of CiBER-seq and its potential for dissecting genetic networks in mammalian cells with superior time resolution.
Chen, W.; Choi, J.; Nathans, J. F.; Agarwal, V.; Martin, B.; Nichols, E.; Leith, A.; Lee, C.; Shendure, J.
Show abstract
Measurements of gene expression and signal transduction activity are conventionally performed with methods that require either the destruction or live imaging of a biological sample within the timeframe of interest. Here we demonstrate an alternative paradigm, termed ENGRAM (ENhancer-driven Genomic Recording of transcriptional Activity in Multiplex), in which the activity and dynamics of multiple transcriptional reporters are stably recorded to DNA. ENGRAM is based on the prime editing-mediated insertion of signal- or enhancer-specific barcodes to a genomically encoded recording unit. We show how this strategy can be used to concurrently genomically record the relative activity of at least hundreds of enhancers with high fidelity, sensitivity and reproducibility. Leveraging synthetic enhancers that are responsive to specific signal transduction pathways, we further demonstrate time- and concentration-dependent genomic recording of Wnt, NF-{kappa}B, and Tet-On activity. Finally, by coupling ENGRAM to sequential genome editing, we show how serially occurring molecular events can potentially be ordered. Looking forward, we envision that multiplex, ENGRAM-based recording of the strength, duration and order of enhancer and signal transduction activities has broad potential for application in functional genomics, developmental biology and neuroscience.
Tran, V.; Papalexi, E.; Schroeder, S.; Kim, G.; Sapre, A.; Pangallo, J.; Sova, A.; Matulich, P.; Kenyon, L.; Sayar, Z.; Koehler, R.; Diaz, D.; Gadkari, A.; Howitz, K.; Nigos, M.; Roco, C. M.; Rosenberg, A. B.
Show abstract
Single cell RNA sequencing (scRNA-seq) has become a core tool for researchers to understand biology. As scRNA-seq has become more ubiquitous, many applications demand higher scalability and sensitivity. Split-pool combinatorial barcoding makes it possible to scale projects to hundreds of samples and millions of cells, overcoming limitations of previous droplet based technologies. However, there is still a need for increased sensitivity for both droplet and combinatorial barcoding based scRNA-seq technologies. To meet this need, here we introduce an updated combinatorial barcoding method for scRNA-seq with dramatically improved sensitivity. To assess performance, we profile a variety of sample types, including cell lines, human peripheral blood mononuclear cells (PBMCs), mouse brain nuclei, and mouse liver nuclei. When compared to the previously best performing approach, we find up to a 2.6-fold increase in unique transcripts detected per cell and up to a 1.8-fold increase in genes detected per cell. These improvements to transcript and gene detection increase the resolution of the resulting data, making it easier to distinguish cell types and states in heterogeneous samples. Split-pool combinatorial barcoding already enables scaling to millions of cells, the ability to perform scRNA-seq on previously fixed and frozen samples, and access to scRNA-seq without the need to purchase specialized lab equipment. Our hope is that by combining these previous advantages with the dramatic improvements to sensitivity presented here, we will elevate the standards and capabilities of scRNA-seq for the broader community.
Payne, A.; Munro, R.; Holmes, N.; Moore, C.; Carlile, M.; Loose, M. W.
Show abstract
Adaptive sampling enables selection of individual DNA molecules from sequencing libraries, a unique property of nanopore sequencing. Here we develop our adaptive sampling tool readfish to become "barcode-aware" enabling selection of different targets within barcoded samples or filtering out individual barcodes. We show that multiple human genomes can be assessed for copy number and structural variation on a single sequencing flow cell using sample specific customised target panels on both GridION and PromethION devices.
Lin, T.-J.; Landry, M. P.
Show abstract
Over 88% of biological research articles use bar graphs, of which 29% have undocumented data distortion mistakes that over- or under-state findings. We developed a framework to quantify data distortion and analyzed bar graphs published across 3387 articles in 15 journals, finding consistent data distortions across journals and common biological data types. To reduce bar graph-induced data distortion, we propose recommendations to improve data visualization literacy and guidelines for effective data visualization.
Lazzarotto, C.; Katta, V.; Li, Y.; Urbina, E.; Lee, G.; Tsai, S. Q.
Show abstract
Base editors (BE) enable programmable conversion of nucleotides in genomic DNA without double-stranded breaks and have substantial promise to become new transformative genome editing medicines. Sensitive and unbiased detection of base editor off-target effects is important for identifying safety risks unique to base editors and translation to human therapeutics, as well as accurate use in life sciences research. However, current methods for understanding the global activities of base editors have limitations in terms of sensitivity or bias. Here we present CHANGE-seq-BE, a novel method to directly assess the off-target profile of base editors that is simultaneously sensitive and unbiased. CHANGE-seq-BE is based on the principle of selective sequencing of adenine base editor modified genomic DNA in vitro, and provides an accessible, rapid, and comprehensive method for identifying genome-wide off-target mutations of base editors.
Gould, S. I.; Sanchez-Rivera, F. J.
Show abstract
Many human diseases have a strong association with diverse types of genetic alterations. These diseases include cancer, in which tumor genomes often harbor a complex spectrum of single-nucleotide alterations and chromosomal rearrangements that can perturb gene function in ways that remain poorly understood. Some cancer-associated genes exhibit a tremendous degree of mutational heterogeneity, which may impact disease initiation, progression, and therapy responses. For example, TP53, the most frequently mutated gene in cancer, shows extensive allelic variation that leads to the generation of altered proteins that can produce functionally distinct phenotypes. Whether distinct variants of TP53 and other genes encode proteins with loss-of-function, gain-of-function, or otherwise neomorphic phenotypes remains both controversial and technically challenging to assess, particularly at the endogenous level. Here, we present a high-throughput prime editing "sensor" strategy to quantitatively assess the functional impact of diverse types of endogenous genetic variants. We used this strategy to screen the largest collection of endogenous cancer-associated TP53 variants assembled to date, identifying both known and novel alleles that impact p53 function in mechanistically diverse ways. Intriguingly, we find that certain types of endogenous TP53 variants, particularly those in the p53 oligomerization domain, display opposite phenotypes in exogenous overexpression systems. These include disease-relevant variants found in humans with cancer predisposition syndromes that encode altered proteins with unique molecular properties. Our results emphasize the physiological importance of gene dosage in shaping native protein stoichiometry and protein-protein interactions, highlight the dangers of using exogenous overexpression systems to interpret pathogenic alleles, and establish a powerful computational and experimental framework for studying diverse types of genetic variants in their endogenous sequence context at scale.