SuPreMo: a computational tool for streamlining in silico perturbation using sequence-based predictive models
Gjoni, K.; Pollard, K. S.
Show abstract
Computationally editing genome sequences is a common bioinformatics task, but current approaches have limitations, such as incompatibility with structural variants, challenges in identifying responsible sequence perturbations, and the need for vcf file inputs and phased data. To address these bottlenecks, we present Sequence Mutator for Predictive Models (SuPreMo), a scalable and comprehensive tool for performing in silico mutagenesis. We then demonstrate how pairs of reference and perturbed sequences can be used with machine learning models to prioritize pathogenic variants or discover new functional sequences. Availability and ImplementationSuPreMo was written in Python, and can be run using only one line of code to generate both sequences and 3D genome disruption scores. The codebase, instructions for installation and use, and tutorials are on the Github page: https://github.com/ketringjoni/SuPreMo/tree/main. Contactkatherine.pollard@gladstone.ucsf.edu Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CLUES2 Companion: Computational pipelines to estimate, visualize, and date selection on multi-locus sites 95%
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 94%
- K2R: Tinted de Bruijn Graphs implementation for efficient read extraction from sequencing datasets 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.