A composite method to infer drug resistance with mixed genomic data
Datta, G.; Hasan, N. A.; Strong, M.; Leach, S. M.
Show abstract
BackgroundThe increasing incidence of drug resistance in tuberculosis and other infectious diseases poses an escalating cause for concern, emphasizing the urgent need to devise robust computational and molecular methods identify drug resistant strains. Although machine learning-based approaches using whole-genome sequence data can facilitate the inference of drug resistance, current implementations do not optimally take advantage of information in public databases and are not robust for small sample sizes and mixed attribute types. ResultsIn this paper we introduce the Composite MetaDistance method, an approach for feature selection and classification of high-dimensional, unbalanced datasets with mixed attribute features from various data sources. We introduce a mixed-attribute, multi-view distance function to calculate distances between samples, with optimal handling of nominal features and different feature views. We also introduce a novel feature set for drug resistance prediction in Mycobacterium tuberculosis, using data from diverse sources. We compare the performance of Composite MetaDistance to multiple machine learning algorithms for Mycobacterium tuberculosis drug resistance prediction for three drugs. Composite MetaDistance consistently outperforms existing algorithms for small sample training sets, and performs as well as other algorithms for training sets with larger sample sizes. ConclusionThe feature set formulation introduced in this paper is utilizes mutational and publicly available information for each gene, and is much richer than ever devised previously. The prediction algorithm, Composite MetaDistance, is sample size agnostic and robust especially given small sample sizes. Proper handling of nominal features improves performance even with a very small number of nominal features. We expect Composite MetaDistance to be even more robust for datasets with a higher percentage of nominal features. The algorithm is application independent and can be used for any mixed attribute dataset.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 94%
- An accurate and interpretable model for antimicrobial resistance in pathogenic Escherichia coli from livestock and companion animal species 93%
- A hierarchical Bayesian latent class mixture model with censorship for detection of linear temporal changes in antibiotic resistance 93%
Similar papers in this journal
- Comparative analysis of machine learning algorithms on the microbial strain-specific AMP prediction 94%
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 93%
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 93%
Similar papers in this journal
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 93%
- DeepSEA: an alignment-free deep learning tool for functional annotation of antimicrobial resistance proteins 93%
- Primary case inference in viral outbreaks through analysis of intra-host variant population 93%
Similar papers in this journal
- Predicting the Trend of SARS-CoV-2 Mutation Frequencies Using Historical Data 93%
- Attention-based approach to predict drug-target interactions across seven target superfamilies 93%
- Demixer: A probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.