Machine learning reveals conserved chromatin patterns determining meiotic recombination sites in plants
Wang, M.; Shilo, S.; Zhou, A.; Zelkowski, M.; Olson, M. A.; Azuri, I.; Shoshani-Hechel, N.; Melamed-Bessudo, C.; Marand, A. P.; Jiang, J.; Schnable, J. C.; Underwood, C. J.; Henderson, I. R.; Sun, Q.; Pillardy, J.; Kianian, P. M. A.; Kianian, S. F.; Chen, C.; Levy, A. A.; Pawlowski, W. P.
Show abstract
Distribution of meiotic recombination events in plants has been associated with local chromatin and DNA characteristics, chromosome landmark proximity, and other features1-7. However, relative importance of these characteristics is unclear and it is unknown if they are sufficient to unambiguously determine recombination landscape8. Here, we analyzed over 40 DNA sequence, chromatin, and chromosome location features of maize and Arabidopsis recombination sites using machine learning9,10. We discovered that a combination of just three features, CG methylation, CHG methylation, and nucleosome occupancy, enabled identification of exact crossover site with 90% accuracy. These results imply redundancy of most recombination site characteristics. Recombination takes place in a small fraction of the genome with chromatin features distinct from those of genome at large. Surprisingly, crossover sites show elevated heterochromatin histone marks despite low DNA methylation. Crossover site features show broad evolutionary conservation, which will enable creating genetic maps in species where conventional mapping is unfeasible.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A highly contiguous genome assembly of Brassica nigra (BB) and revised nomenclature for the pseudochromosomes 92%
- CNCC: An analysis tool to determine genome-wide DNA break end structure at single-nucleotide resolution 92%
- To what extent gene connectivity within co-expression network matters for phenotype prediction? 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.