Data coherence over data volume drives generalisable genome-based prediction of microbial carbon utilisation
Kishore, D.; Ranjan, P.; Neely, C.; Cashman, M.; Riehl, W.; Joachimiak, M. P.; Edirisinghe, J. N.; Faria, J. P.; Cohen, M. B.; Sakkaff, Z.; Weisenhorn, P.; Pelletier, D. A.; Doktycz, M. J.; Cottingham, R. W.; Henry, C. S.; Arkin, A. P.; Dehal, P. S.
Show abstract
Microbial carbon utilisation is a foundational ecological phenotype that remains difficult to predict from genomes despite well-characterised pathways. Machine-learning models generalise poorly across datasets, a failure usually attributed to training-set size and taxonomic bias. To test this, we integrated binary growth phenotypes for 819 strains across 240 carbon sources from four datasets. Balanced accuracy fell from 0.86 within datasets to 0.62 across them, and testing on close relatives recovered only 0.03 of that drop, so mechanistically inconsistent genotype-phenotype relationships drove models to dataset-correlated shortcuts. Restricting training to concordant samples (measured growth matched their annotated pathway) doubled the carbon sources recovering known pathway genes across datasets (6 to 12 of 15), whereas matched random subsets did not. Adding over 8000 literature-curated BacDive genomes to the training set did not improve cross-dataset performance more than the smaller concordant set, suggesting coherence matters more than volume. Because such filtering requires a mechanistic predictor, we tested a mechanism-free alternative combining phylogenetic agreement and experimental labels, which recovered part of the gain but not the recall advantage. Concordance-trained models were bounded specialists, rescuing mechanistic false negatives twice as often as false positives (44% versus 19%), mostly metabolic generalists. To locate those bounds, model confidence defined an applicability domain, and prioritising low-confidence genomes for training improved cross-dataset accuracy more than random or diversity-based sampling, especially for the weaker phenotypes. This recasts generalisation in biological machine learning as a problem of label-mechanism agreement and applicability-domain definition, alongside data volume and algorithm choice.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integrating theory and machine learning to reveal determinants of plasmid copy number 94%
- Systematic evaluation of metatranscriptomic differential gene expression in silico, in vitro, and in vivo enables elucidation of inter-species cross-feeding 94%
- There is no evidence of a universal genetic boundary among microbial species 94%
Similar papers in this journal
- Metabolic rules of microbial community assembly 94%
- APOLLO: A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites 93%
- Evolution in microbial microcosms is highly parallel regardless of the presence of interacting species 93%
Similar papers in this journal
Similar papers in this journal
- Quantifying biosynthetic network robustness across the human oral microbiome 94%
- Automated genome mining predicts structural diversity and taxonomic distribution of peptide metallophores across bacteria 92%
- Evolutionary stability of collateral sensitivity to antibiotics in the model pathogen Pseudomonas aeruginosa 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.