Few-shot in-context learning with large language models for antibody characterization
FUNG, S.-H.; Zhang, Z.; Wang, R.; Miao, C.; Wong, B. S.-H.; Li, K. Y.; Hong, C.; Zhou, J.; Yip, K. Y.; Tsui, S. K.-W.; Cao, Q.
Show abstract
Large language models (LLMs) can learn new tasks by in-context learning (ICL), but it is unknown whether this ability reliably transfers to biological sequence classification. Here, we systematically evaluate how demonstration selection, shot count, and prompting strategies affect performance across 20 general-purpose LLMs. Using antibody characterization as a representative test case, we compare zero-shot, few-shot, and chain-of-thought (CoT) ICL on three classification tasks: humanness, antigen specificity, and isotype. Our results reveal a clear performance hierarchy: while zero-shot prompting performs near chance, few-shot prompting with randomly selected demonstrations improves performance, showing that LLMs can perform ICL using biological sequences from minimal supervision. However, matching protein-language model (pLM)-based classifier accuracy is only achieved when using label-diverse demonstrations drawn from antibodies similar to the query sequence. To leverage this insight, we introduce Sim-ICL, a framework that automatically retrieves such demonstrations. Using only 32-shot prompting, Sim-ICL achieves performance competitive with pLM-based classifiers, matching or outperforming them in two of the three tasks. Furthermore, reasoning-oriented prompts yield marginal gains and often produce fluent but biologically incorrect rationales, suggesting that current CoT explanations function as after-the-fact rationalizations rather than capturing mechanistic determinants of antibody properties. From these experiments, we derive practical design principles for ICL on biological sequences: use similarity-based, label-diverse demonstrations and modest shot counts, and treat reasoning prompts primarily as post hoc narratives rather than drivers of performance. Sim-ICL implements these principles in a streamlined, prompt-based framework for antibody sequence classification and, in principle, could be adapted to other biological sequence tasks.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting the antigenic evolution of SARS-COV-2 with deep learning 96%
- Triple-effect correction for Cell Painting data with contrastive and domain-adversarial learning 96%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 95%
Similar papers in this journal
Similar papers in this journal
- Deep autoregressive generative models capture the intrinsics embedded in T-cell receptor repertoires 96%
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 95%
- HyGAnno: Hybrid graph neural network-based cell type annotation for single-cell ATAC sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.