Back

Machine Learning Unravels Inherent Structural Patterns in Escherichia Coli Hi-C Matrices and Predicts DNA Dynamics

Bera, P.; Mondal, J.

2023-12-20 biophysics
10.1101/2023.12.20.572497 bioRxiv
Show abstract

The large dimension of the Hi-C-derived chromosomal contact map, even for a bacterial cell, presents challenges in extracting meaningful information related to its complex organization. Here we first demonstrate that a machine-learnt (ML) low-dimensional embedding of a recently reported Hi-C interaction map of archetypal bacteria E. Coli can decode crucial underlying structural pattern. In particular, a three-dimensional latent space representation of (928x928) dimensional Hi-C map, derived from an unsupervised artificial neural network, automatically detects a set of spatially distinct domains that show close correspondences with six macro-domains (MDs) that were earlier proposed across E. Coli genome via recombination assay-based experiments. Subsequently, we develop a supervised random-forest regression model by machine-learning intricate relationship between large array of Hi-C-derived chromosomal contact probabilities and diffusive dynamics of each individual chromosomal gene. The resultant ML model dictates that a minimal subset of important chromosomal contact pairs (only 30 %) out of full Hi-C map is sufficient for optimal reconstruction of the heterogenous, coordinate-dependent sub-diffusive motions of chromosomal loci. Specifically the Ori MD was predicted to exhibit most substantial contribution in chromosomal dynamics among all MDs. Finally, the ML models, trained on wild-type E. Coli was tested for its predictive capabilities on mutant bacterial strains, shedding light on the structural and dynamic nuances of {Delta}MatP30MM and {Delta}MukBEF22MM chromosomes. Overall our results illuminate the power of ML techniques in unraveling the complex relationship between structure and dynamics of bacterial chromosomal loci, promising meaningful connections between our ML-derived insights and real-world biological phenomena.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.