A fast and objective hidden Markov modeling for accurate analysis of biophysical data with numerous states
Liu, H.; Shima, T.
Show abstract
The hidden Markov model (HMM) is widely used to analyze biophysical chronological data with discrete states, such as binding/detachment of biomolecules, protein/nucleotide conformational changes and step-like movement of single proteins. Despite its usefulness, classical HMM fitting has practical drawbacks that it requires the determination of the number of hidden states and fine initialization of many parameters before fitting. To overcome these drawbacks, several HMM pre-analyses have been reported, but do not provide enough accuracy when data have unknown kinetics and/or low signal-to-noise ratio. Therefore, in many cases, HMM fitting needs trial-and-error manual process that can impair the objectivity of the analysis. Moreover, for data composed of numerous hidden states, such as stepping data of cytoskeletal motors, there has been difficulty in HMM analysis because the large number of parameters were hardly properly initialized. Here, by combining a statistical step-finding method and the Gaussian mixture model clustering, we developed a new algorithm for more objective HMM analysis. Our algorithm can execute accurate state number estimation and parameter optimization with fully automated way. Simulation analysis demonstrated that our algorithm accurately fit both fast- and slow-transition trajectories. Compared with the previous method, the speed of our algorithm was 10-20 times faster for standard size data. Our algorithm also showed the accurate fit of the simulated motor-stepping data with more than 10 transition states, suggesting the applicability of the method to the data with numerous states. Furthermore, the algorithm is flexible enough to cope with cases where some kinetics are known in advance. Some available prior information, such as the dwell time of each state, can be integrated into the algorithm via two user-tunable parameters. In summary, our method enables fast, accurate and objective HMM analysis, and broadens the application range of HMM fitting that can provide more accurate interpretation of a wide variety of biophysical data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cadherin clusters stabilized by a combination of specific and nonspecific Cis-Interactions 91%
- Wavenumber-dependent transmission of subthreshold waves on electrical synapses network model of Caenorhabditis elegans 91%
- Deciphering anomalous heterogeneous intracellular transport with neural networks 91%
Similar papers in this journal
- Storm: Incorporating transient dynamics to infer the RNA velocity with metabolic labeling information 93%
- Beam search decoder for enhancing sequence decoding speed in single-molecule peptide sequencing data 93%
- Assessing the Performance of Methods for Cell Clustering from Single-cell DNA Sequencing Data 92%
Similar papers in this journal
- RevGraphVAMP: A protein molecular simulation analysis model combining graph convolutional neural networks and physical constraints 92%
- LoopSage: An Energy-Based Monte Carlo approach for the Loop Extrusion Modelling of Chromatin 92%
- Penguin: A Tool for Predicting Pseudouridine Sites in Direct RNA Nanopore Sequencing Data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.