Reliable Evaluation and Learning in Multi-input Biological Association Prediction
Ahmadian Moghadam, S.; Montazeri, H.
Show abstract
Multi-input association prediction is central to many key problems in computational biology, spanning tasks from drug-target and protein-protein interactions to higher-order challenges such as drug synergy modeling and MHC-peptide-TCR binding. Yet widely used benchmarks often overestimate performance by enabling models to exploit degree ratio shortcuts, while alternative out-of-distribution splits are overly restrictive and impractical. Here we introduce an entity-balanced evaluation framework that systematically neutralizes shortcut signals by balancing positive and negative associations at the entity level. This enables fairer assessments that reflect genuine relational learning and extend naturally from pairwise to multi-entity problems. We further present UnbiasNet, a model-agnostic training strategy that cycles through diverse entity-balanced sub-training sets, removing access to degree ratio bias and enhancing robustness. Applied to drug-target interaction and drug synergy prediction, our framework reveals the extent of shortcut reliance in existing methods while enabling consistent identification of meaningful biological associations, thereby setting a rigorous foundation for future methodological progress.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 96%
- Hierarchical affinity landscape navigation through learning a shared pocket-ligand space 95%
- Tokenized and Continuous Embedding Compressions of Protein Sequence and Structure 94%
Similar papers in this journal
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 95%
- JIND: Joint Integration and Discrimination for Automated Single-Cell Annotation 95%
- KG-Bench: Benchmarking Graph Neural Network Algorithms for Drug Repurposing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.