OptiK: An Entropy-Driven Framework for Optimal k-mer Size Selection for Bacterial Genomics
Gutierrez, A.
Show abstract
K-mer-based approaches have become fundamental (Zielezinski et al., 2017) to modern computational genomics, underpinning tools for genome assembly, metagenomic classification, variant calling, and phylogenetic analysis. Despite their ubiquity, selecting an appropriate k-mer size (k) is often made arbitrarily or heuristically, with little consideration for the underlying signal quality relative to a given dataset. Here, I introduce OptiK, a novel alignment-free tool that evaluates the information richness of k-mer encodings across a range of k values to identify the optimal k for comparative analysis. OptiK operates by constructing k-mer frequency matrices from genome collections, reducing their dimensionality via truncated singular value decomposition (SVD), and evaluating clustering structure through unsupervised metrics including the Silhouette coefficient, Calinski-Harabasz index, and Davies-Bouldin index. We validate OptiK on a curated dataset of 1044 Helicobacter pylori genomes with well-characterized population structure. OptiK robustly identifies k = 8 as the optimal k-mer size, yielding latent structures in UMAP space that align with fineSTRUCTURE-defined subpopulations without relying on prior labels or reference alignments. These results demonstrate that OptiK provides a reproducible, alignment-free strategy for optimizing k-mer resolution in bacterial comparative genomics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AdDeam: A Fast and Scalable Tool for Estimating and Clustering Reference-Level Damage Profiles 95%
- The Naive Bayes Classifier++ for Metagenomic Taxonomic Classification -- Query Evaluation 95%
- Themisto: a scalable colored k-mer index for sensitive pseudoalignment against hundreds of thousands of bacterial genomes 95%
Similar papers in this journal
- K2R: Tinted de Bruijn Graphs implementation for efficient read extraction from sequencing datasets 94%
- MerCat2: a versatile k-mer counter and diversity estimator for database-independent property analysis obtained from omics data 94%
- baseLess: Lightweight detection of sequences in raw MinION data 94%
Similar papers in this journal
- Sketching and sampling approaches for fast and accurate long read classification 94%
- HapSolo: An optimization approach for removing secondary haplotigs during diploid genome assembly and scaffolding. 94%
- NucBreak: Location of structural errors in a genome assembly by using paired-end Illumina reads 93%
Similar papers in this journal
Similar papers in this journal
- PyOrthoANI, PyFastANI, and Pyskani: a suite of Python libraries for computation of average nucleotide identity 94%
- Metagenomics-Toolkit: The Flexible and Efficient Cloud-Based Metagenomics Workflow featuring Machine Learning-Enabled Resource Allocation 94%
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.