Active Learning for Budget-Constrained TCR--pMHC Wet-Lab Validation
Mazur, K.; Piotrowska, M.; Kowalski, J.
Show abstract
Wet-lab validation of TCR-pMHC binding hypotheses is the rate-limiting step in T-cell therapy discovery: a single binding assay round can cost thousands of dollars and weeks of turnaround time, yet computational models generate thousands of candidate pairs per run. We frame this as a pool-based active learning problem: given a fixed annotation budget B, which unlabeled pairs should be sent to the assay to maximally improve a predictive model that will guide the next screening round? We introduce UDAL (Uncertainty-Diversity Active Learning), a batch acquisition strategy that combines BALD-based uncertainty estimation via MC Dropout with greedy core-set diversity selection in the encoder feature space. Evaluated on a curated VDJdb-IEDB benchmark under epitope-held-out and distance-aware protocols, UDAL achieves AUPRC 0.487 with only 5,000 queried labels--matching the performance of a model trained on 3 x more randomly sampled labels. At a budget of 2,000 labels, UDAL improves AUPRC by 16.7% over random acquisition, translating directly to fewer wasted assay slots. These results demonstrate that principled active query strategies can substantially reduce the wet-lab cost of building reliable TCR specificity models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EPIC-TRACE: predicting TCR binding to unseen epitopes using attention and contextualized embeddings 95%
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 94%
- Learning Context-aware Structural Representations to Predict Antigen and Antibody Binding Interfaces 94%
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- Benchmarking Uncertainty Quantification for Protein Engineering 95%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 94%
Similar papers in this journal
- Designing meaningful continuous representations of T cell receptor sequences with deep generative models 94%
- Sparse Epistatic Regularization of Deep Neural Networks for Inferring Fitness Functions 94%
- Machine Learning Optimization of Candidate Antibodies Yields Highly Diverse Sub-nanomolar Affinity Antibody Libraries 93%
Similar papers in this journal
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 94%
- TEINet: a deep learning framework for prediction of TCR-epitope binding specificity 94%
- Synthetic observations from deep generative models and binary omics data with limited sample size 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.