Machine learning to classify mutational hotspots from molecular dynamics simulations
Menzies, G. E.; Davies, J.
Show abstract
Benzo[a]pyrene, a notorious DNA-damaging carcinogen, belongs to the family of polycyclic aromatic hydrocarbons commonly found in tobacco smoke. Surprisingly, nucleotide excision repair (NER) machinery exhibits inefficiency in recognising specific bulky DNA adducts including Benzo[a]pyrene Diol-Epoxide (BPDE), a Benzo[a]pyrene metabolite. While sequence context is emerging as the leading factor linking the inadequate NER response to BPDE adducts, the precise structural attributes governing these disparities remain inadequately understood. We therefore combined the domains of molecular dynamics and machine learning to conduct a comprehensive assessment of helical distortion caused by BPDE-Guanine adducts in multiple gene contexts. Specifically, we implemented a dual approach involving a random forest classification-based analysis and subsequent feature selection to identify precise topological features that may distinguish adduct sites of variable repair capacity. Our models were trained using helical data extracted from duplexes representing both BPDE hotspot and non-hotspot sites within the TP53 gene, then applied to sites within TP53, cII, and lacZ genes. We show our optimised model consistently achieved exceptional performance, with accuracy, precision, and f1 scores exceeding 91%. Our feature selection approach uncovered that discernible variance in regional base pair rotation played a pivotal role in informing the decisions of our model. Notably, these disparities were highly conserved among TP53 and lacZ duplexes and appeared to be influenced by the regional GC content. As such, our findings suggest that there are indeed conserved topological features distinguishing hotspots and non-hotpot sites, highlighting regional GC content as a potential biomarker for mutation. Author SummaryAlthough much is known about DNA repair processes, we are still lacking some fundamental understanding relating to DNA sequence and mutation rates, specifically why some sequences mutate at a higher rate or are repaired less than others. We believe that by using a combination of Molecular Simulation and Machine Learning (ML) we can measure which structural features are present in sequences which mutate at higher rates in cancer gene and lab-based test assays frequently used to investigate toxicology. Here we have run Molecular Dynamics on five sets of DNA sequences with and without a carcinogen found in cigarette smoke to allow us to study the mutation event that would need to be repaired. We have measured their helical and base stacking properties. We have used ML to successfully differentiate between low and high mutating sequences using this model allowing us to begin to elucidate the structural features these groups have in common. We believe this method could have wide reaching uses, it could be applied to any gene context and mutation event and indeed the knowledge of the structural features which are best repaired gives us insight into the biophysics of DNA repair adding knowledge to the drug design pipeline.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Solvation of the E. coli CheY Phosphorylation SiteMapped by XFMS 92%
- Single nucleotide polymorphism induces divergent dynamic patterns in CYP3A5: a microsecond scale biomolecular simulation of variants identified in Sub-Saharan African populations 91%
- Variable inhibition of unwinding rates of DNA catalyzed by the SARS-Cov-2 (COV19) helicase nsp13 by structurally distinct single DNA lesions. 91%
Similar papers in this journal
- The DNA Damage-Sensing NER Repair Factor XPC-RAD23B Does Not Recognize Bulky DNA Lesions with a Missing Nucleotide Opposite the Lesion. 94%
- Overexpression of the WWE domain of RNF146 modulates poly-(ADP)-ribose dynamics at sites of DNA damage 89%
- Corruption of DNA End-Joining in Mammalian Chromosomes by Progerin Expression as Revealed by a Model Cell Culture System 89%
Similar papers in this journal
- Quantum biological insights into CRISPR-Cas9 sgRNA efficiency from explainable-AI driven feature engineering 93%
- RaptRanker: in silico RNA aptamer selection from HT-SELEX experiment based on local sequence and structure information 92%
- Discovery and Development of Novel DNA-PK Inhibitors by Targeting the unique Ku-DNA Interaction 92%
Similar papers in this journal
- Extraction and high-throughput sequencing of oak heartwood DNA: assessing the feasibility of genome-wide DNA methylation profiling 92%
- PFRED: A computational platform for siRNA and antisense oligonucleotides design. 92%
- Single molecule studies characterize the kinetic mechanism of tetrameric p53 binding to different native response elements 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.