Back

Data coherence over data volume drives generalisable genome-based prediction of microbial carbon utilisation

Kishore, D.; Ranjan, P.; Neely, C.; Cashman, M.; Riehl, W.; Joachimiak, M. P.; Edirisinghe, J. N.; Faria, J. P.; Cohen, M. B.; Sakkaff, Z.; Weisenhorn, P.; Pelletier, D. A.; Doktycz, M. J.; Cottingham, R. W.; Henry, C. S.; Arkin, A. P.; Dehal, P. S.

2026-08-12 bioinformatics
10.64898/2026.08.06.743333 bioRxiv
Show abstract

Microbial carbon utilisation is a foundational ecological phenotype that remains difficult to predict from genomes despite well-characterised pathways. Machine-learning models generalise poorly across datasets, a failure usually attributed to training-set size and taxonomic bias. To test this, we integrated binary growth phenotypes for 819 strains across 240 carbon sources from four datasets. Balanced accuracy fell from 0.86 within datasets to 0.62 across them, and testing on close relatives recovered only 0.03 of that drop, so mechanistically inconsistent genotype-phenotype relationships drove models to dataset-correlated shortcuts. Restricting training to concordant samples (measured growth matched their annotated pathway) doubled the carbon sources recovering known pathway genes across datasets (6 to 12 of 15), whereas matched random subsets did not. Adding over 8000 literature-curated BacDive genomes to the training set did not improve cross-dataset performance more than the smaller concordant set, suggesting coherence matters more than volume. Because such filtering requires a mechanistic predictor, we tested a mechanism-free alternative combining phylogenetic agreement and experimental labels, which recovered part of the gain but not the recall advantage. Concordance-trained models were bounded specialists, rescuing mechanistic false negatives twice as often as false positives (44% versus 19%), mostly metabolic generalists. To locate those bounds, model confidence defined an applicability domain, and prioritising low-confidence genomes for training improved cross-dataset accuracy more than random or diversity-based sampling, especially for the weaker phenotypes. This recasts generalisation in biological machine learning as a problem of label-mechanism agreement and applicability-domain definition, alongside data volume and algorithm choice.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.