CentroFinder: accurate de novo identification of centromeres in fungal genomes
Salimi, S.; Colson, S.; Renfro, M.; Ma, L.-J.; Rahnama, M.
Show abstract
MotivationCentromeres are essential chromosomal loci, yet their computational identification remains challenging due to rapid sequence evolution, high repeat content, and the absence of conserved defining motifs. This challenge is particularly pronounced in fungi, where centromere architectures vary widely in size, sequence composition, and chromatin organization, limiting the effectiveness of single-feature or motif-based prediction approaches. ResultsWe present CentroFinder, a fungal-specific computational framework for de novo centromere prediction from long-read sequencing-based genome assemblies. CentroFinder integrates multiple genomic and long-read-derived features into a weighted scoring model to identify loci where centromere-associated signals converge. Benchmarking against experimentally mapped centromeres in Cryptococcus deuterogattii, Magnaporthe oryzae, and Neurospora crassa demonstrates that CentroFinder consistently predicts a single centromeric region per chromosome, fully nested within CENP-A-defined domains despite substantial diversity in centromere size, sequence composition, and chromatin context. Availability and ImplementationCentroFinder is freely available as open-source software at https://github.com/RahnamaLab/CentroFinder. The pipeline is designed for high-performance computing environments and leverages features derived from long-read sequencing data.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ancestral and recent bursts of transposition shaped the massive genomes of plant pathogenic rust fungi 92%
- Benchmarking long-read variant calling in diploid and polyploid genomes: insights from human and plants 92%
- Evolutionary genomics reveals variation in structure and genetic content implicated in virulence and lifestyle in the genus Gaeumannomyces 91%
Similar papers in this journal
- PoMeLo: a systematic computational approach to predicting metabolic loss in pathogen genomes 91%
- SpectralTAD: an R package for defining a hierarchy of Topologically Associated Domains using spectral clustering 91%
- GenErode: a bioinformatics pipeline to investigate genome erosion in endangered and extinct species 90%
Similar papers in this journal
- Higher order repeat structures reflect diverging evolutionary paths in maize centromeres and knobs 94%
- Chromosome-level quality scaffolding of brown algal genomes using InstaGRAAL, a proximity ligation-based scaffolder 93%
- Automated assembly scaffolding elevates a new tomato system for high-throughput genome editing 92%
Similar papers in this journal
- Homoeologous gene expression and co-expression network analyses and evolutionary inference in allopolyploids 91%
- binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets 91%
- Comprehensive benchmarking of software for mapping whole genome bisulfite data: from read alignment to DNA methylation analysis 90%
Similar papers in this journal
- The genome of the oomycete Peronosclerospora sorghi, a cosmopolitan pathogen of maize and sorghum, is inflated with dispersed pseudogenes 93%
- Host adaptation and genome evolution of the broad host range fungal rust pathogen, Austropuccinia psidii 93%
- Genome Dynamics and Chromosome Structural Variations in Histoplasma ohiense, a fungal pathogen of humans 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.