Prediction of representative phenotypes using multi-output subset selection
Forchielli, E.; Wang, T.; Thommes, M.; Paschalidis, I. C.; Segre, D.
Show abstract
The interpretation of complex biological datasets requires the identification of representative variables that describe the data without critical information loss. This is particularly important in the analysis of large phenotypic datasets ("phenomics"). We introduce Multi-Attribute Subset Selection (MASS), an algorithm which separates a matrix of phenotypes (e.g., yield across microbial species and environmental conditions) into predictor and response sets of conditions. Using mixed integer linear programming, MASS expresses the response conditions as a linear combination of the predictor conditions, while simultaneously searching for the optimally descriptive set of predictors. We applied the algorithm to three microbial datasets and identified environmental conditions that predict phenotypes under other conditions, providing biologically interpretable axes for strain discrimination. MASS could be used to reduce the number of experiments needed to identify species or to map their metabolic capabilities. The generality of the algorithm allows addressing subset selection problems in areas beyond biology.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Guiding the Refinement of Biochemical Knowledgebases with Ensembles of Metabolic Networks and Machine Learning 96%
- Inferring protein sequence-function relationships with large-scale positive-unlabeled learning 91%
- Pitfalls of genotyping microbial communities with rapidly growing genome collections 91%
Similar papers in this journal
Similar papers in this journal
- Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions 95%
- Genome-scale metabolic modelling when changes in environmental conditions affect biomass composition 94%
- A pipeline for the reconstruction and evaluation of context-specific human metabolic models at a large-scale 94%
Similar papers in this journal
- Calcium starvation leads to strain-specific gene regulation of lipid and carotenoid production in Mucor Circinelloides 92%
- Yeast population dynamics in Brazilian bioethanol production 92%
- Genomic and Transcriptomic Characterization of Carbohydrate-Active Enzymes in the Anaerobic Fungus Neocallimastix cameroonii var. constans 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.