Protein Language Models are Accidental Taxonomists
Hallee, L. P.; Peleg, T.; Rafailidis, N.; Gleghorn, J. P.
Show abstract
Protein-protein interactions (PPIs) are fundamental to nearly all biological processes, yet their experimental characterization remains costly and time-consuming. While computational methods, particularly those using protein language models (pLMs), offer higher-throughput solutions, they often report unexpectedly high performance on multi-species datasets. Here, we introduce the accidental taxonomist hypothesis, proposing that neural networks can exploit the phylogenetic distances across labels in protein datasets rather than genuine interaction features. We show that in PPI datasets with random negative sampling, protein pairs for real PPIs are almost exclusively from the same species, while negatives almost always originate from different species. We then demonstrate that pLM embeddings can be used to accurately distinguish whether two proteins share a taxonomic origin, allowing models to "cheat" by learning phylogeny instead of genuine PPI features. By employing a strategic sampling strategy that restricts negative examples to protein pairs from the same species, we reveal a marked drop in model performance, confirming our hypothesis. Compellingly, these strategically trained models still outperform single-species models, suggesting that multi-species data can improve performance if carefully curated. These findings suggest that accidental taxonomist behavior is a particularly influential confounder for PPI, and it is also broadly applicable to any supervised-learning protein dataset.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 98%
- Rossmann-toolbox: a deep learning-based protocol for the prediction and design of cofactor specificity in Rossmann-fold proteins 95%
- A simple workflow to identify novel Small Linear Motif (SLiM)-mediated interactions with AlphaFold 94%
Similar papers in this journal
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 95%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 94%
- Discovering differential genome sequence activity with interpretable and efficient deep learning 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.