Leveraging FracMinHash Containment for Genomic dN/dS
Rodriguez, J. S.; Hera, M. R.; Koslicki, D.
Show abstract
Increasing availability of genomic data demands algorithmic approaches that can efficiently and accurately conduct downstream genomic analyses. These analyses, such as evaluating selection pressures within and across genomes, can reveal developmental and environmental pressures. One such commonly used metric to measure evolutionary pressures is based on the ratio of non-synonymous and synonomous substitution rates, dN/dS. Conventionally, the dN/dS ratio is used to infer selection pressures employing alignments to estimate total non-synonymous and synonymous substitution rates along protein-coding genes. However, this process can be time consuming and not scalable for larger datasets. Recently, a fast, approximate similarity measure, FracMinHash containment, was introduced and related to average nucleotide identity. In this work, we show how FracMinHash containment can be used to quickly estimate dN/dS enabling alignment-free estimations at a genomic level. Through simulated and real world experiments, our results indicate that employing FracMinHash containment to estimate dN/dS is scalable, enabling pairwise dN/dS estimations for 85,205 genomes within 5 hours. Furthermore, our approach is comparable to traditional dN/dS methods, representing sequences subject to positive and negative selection across various mutation rates. Moreover, we used this model to evaluate signatures of selection between Archaeal and Bacterial genomes, identifying a previously unreported metabolic island between Methanobrevibacter sp. RGIG2411 and Candidatus Saccharibacteria bacterium RGIG2249. We present, FracMinHash dN/dS, a novel alignment-free approach for estimating dN/dS at a genome level that is accurate and scalable beyond gene-level estimations while demonstrating comparability to conventional alignment-based dN/dS methods. Leveraging the alignment-free similarity estimation, FracMinHash containment, pairwise dN/dS estimations are facilitated within milliseconds, making it suitable for large-scale evolutionary analyses across diverse taxa. It supports comparative genomics, evolutionary inference, and functional interpretation across both synthetic, and complex biological datasets. Availability and implementationA version of the implementation is available at https://github.com/KoslickiLab/dnds-using-fmh.git. The reproduction of figures, data, and analysis can be found at https://github.com/KoslickiLab/dnds-using-fmh_reproducibles.git. Contactdmk333@psu.edu Supplementary informationSupplementary data are available at PLOS Computational Biology online. Author summaryUnderstanding how evolution shapes genomes helps us learn about the pressures organisms face in their environments. Scientists traditionally measure this by comparing genetic changes that alter proteins versus those that dont, a ratio that reveals whether natural selection is preserving or changing genes. However, this conventional approach requires computationally intensive sequence alignments, making it impractical for analyzing the massive genomic datasets now available. We developed a faster, alignment-free method to estimate evolutionary pressure across entire genomes. Our approach uses a computational technique called FracMinHash that compresses genomic information while preserving meaningful patterns. We tested our method on both simulated and real-world data, including over 85,000 microbial genomes, completing the analysis in just five hours whereas traditional methods would take days or weeks for the same analysis. The results were comparable to traditional methods and correctly identified genes under different types of selection. Using this approach, we discovered a previously unreported shared genetic region between an archaeal and bacterial species from the goat gut microbiome, suggesting ancient gene transfer between these distant branches of life. Our method makes large-scale evolutionary analysis practical for diverse applications, from tracking microbial strains to understanding adaptation in complex microbial communities, potentially accelerating discoveries in comparative genomics and evolutionary biology.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Read-SpaM: assembly-free and alignment-free comparison of bacterial genomes with low sequencing coverage 96%
- Statistical Analysis of Variability in TnSeq Data Across Conditions Using Zero-Inflated Negative Binomial Regression 95%
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.