Estimating amino acid substitution models from genome datasets: A simulation study on the performance of estimated models
Nguyen Huy, T.; Dang, C. C.; Sy Vinh, L.
Show abstract
Estimating amino acid substitution models is a crucial task in bioinformatics. The maximum likelihood (ML) approach has been proposed to estimate amino acid substitution models from large datasets. The quality of newly estimated models is normally assessed by comparing with the existing models in building ML trees. Two important questions remained are the correlation of the estimated models with the true models and the required size of the training datasets to estimate reliable models. In this paper, we performed a simulation study to answer these two questions based on the simulated data. We simulated genome datasets with different number of genes/alignments based on predefined models (called true models) and predefined trees (called true trees). The simulated datasets were used to estimate amino acid substitution model using the ML estimation method. Our experiment showed that models estimated by the ML methods from simulated datasets with more than 100 genes have high correlations with the true models. The estimated models performed well in building ML trees in comparison with the true models. The results suggest that amino acid substitution models estimated by the ML methods from large genome datasets might play as reliable tool for analyzing amino acid sequences.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Building alternative consensus trees and supertrees using k-means and Robinson and Foulds distance 96%
- Accurate and Efficient Cell Lineage Tree Inference from Noisy Single Cell Data: the Maximum Likelihood Perfect Phylogeny Approach 94%
- Maximum Likelihood Reconstruction of Ancestral Networks by Integer Linear Programming 93%
Similar papers in this journal
- PsiPartition: Improved Site Partitioning for Genomic Data by Parameterized Sorting Indices and Bayesian Optimization 94%
- Deformity Index: A semi-reference quality metric of phylogenetic trees based on their clades 92%
- Extant Sequence Reconstruction: The accuracy of ancestral sequence reconstructions evaluated by extant sequence cross-validation 92%
Similar papers in this journal
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 94%
- A Phylogenetic Approach to Inferring the Order in Which Mutations Arise during Cancer Progression 94%
- On the automatic annotation of gene functions using observational data and phylogenetic trees 93%
Similar papers in this journal
- Inference of Phylogenetic Networks from Sequence Data using Composite Likelihood 95%
- Assessing Confidence in Root Placement on Phylogenies: An Empirical Study Using Non-Reversible Models for Mammals 94%
- Impact of Ghost Introgression on Coalescent-based Species Tree Inference and Estimation of Divergence Time 94%
Similar papers in this journal
- Evaluating probabilistic programming and fast variational Bayesian inference in phylogenetics 93%
- Genomic comparison of non-photosynthetic plants from the family Balanophoraceae with their photosynthetic relatives. 91%
- A variable-rate quantitative trait evolution model using penalized-likelihood 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.