A fast, general synteny detection engine
Ahrens, J. B.; Wade, K. J.; Pollock, D. D.
Show abstract
The increasingly widespread availability of genomic data has created a growing need for fast, sensitive and scalable comparative analysis methods. A key aspect of comparative genomic analysis is the study of synteny, co-localized gene clusters shared among genomes due to descent from common ancestors. Synteny can provide unique insight into the origin, function, and evolution of genome architectures, but methods to identify syntenic patterns in genomic datasets are often inflexible and slow, and use diverse definitions of what counts as likely synteny. Moreover, the reliable identification of putatively syntenic regions (i.e., whether they are truly indicative of homology) with different lengths and signal to noise ratios can be difficult to quantify. Here, we present Mology, a fast, flexible, alignment-free, nonparametric method to detect regions of syntenic elements among genomes or other datasets. The core algorithm operates on consecutive, rank-ordered elements, which could be genes, operons, motifs, sequence fragments, or any other orderable element. It is agnostic to the physical distance between distinct elements and also to directionality and order within syntenic regions, although such considerations can be addressed post hoc. We describe the underlying statistical theory behind our analysis method, and employ a Monte Carlo approach to estimate the false positive rate and positive predictive values for putative syntenic regions. We also evaluate how varying amounts of noise affect recovery of true syntenic regions among Saccharomycetaceae yeast genomes with up to ~100 million years of divergence. We discuss different strategies for recursive application of our method on syntenic regions with sparser signal than considered here, as well as the general applicability of the core algorithm.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ERC 2.0 - evolutionary rate covariation update provides more powerful inference of functional interactions across large phylogenies 96%
- Assessing and mitigating privacy risk of sparse, noisy genotypes by local alignment to haplotype databases 94%
- ScatTR: Estimating the Size of Long Tandem Repeat Expansions from Short-Reads 94%
Similar papers in this journal
- COATi: statistical pairwise alignment of protein-coding sequences 95%
- Phylogenetic Permulations: a statistically rigorous approach to measure confidence in associations between phenotypes and genetic elements in a phylogenetic context 95%
- Ongoing recombination in SARS-CoV-2 revealed through genealogical reconstruction 95%
Similar papers in this journal
- pSONIC: Ploidy-aware Syntenic Orthologous Networks Identified via Collinearity 95%
- Chromonomer: a tool set for repairing and enhancing assembled genomes through integration of genetic maps and conserved synteny 94%
- Testcrosses are an efficient strategy for identifying cis regulatory variation: Bayesian analysis of allele specific expression (BASE) 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.