EvoStructCLIP: A Mutation-Centered Multimodal Embedding Model for CAGI7 Variant Effect Prediction
Chung, K.; Lee, J.; Kim, Y.; Lee, J.; Park, J.; Lee, H.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWWe present EvoStructCLIP, a mutation-centered multimodal embedding model that integrates local 3D structural windows and evolutionary constraints to predict missense variant effects. EvoStructCLIP combines two encoders: a structure voxel encoder derived from AlphaFold residue neighborhoods and an MSA-based evolutionary encoder. It aligns the modalities through CLIP-style contrastive learning, with FuseMix regularization and an auxiliary pathogenicity loss trained on 153,787 ClinVar variants. Evaluations using lightweight regressors demonstrate that EvoStructCLIP embeddings capture highly transferable predictive signals across diverse phenotypes, including gene-specific functional readouts of BRCA1, KCNQ4, and PTEN/TPMT. This transferability is further supported in the CAGI7 blind competition setting, where models generalized to predicting different gene-specific readouts for BARD1, FGFR, and TSC2 without target-specific retraining and achieved competitive performance across heterogeneous biological tasks.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 96%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 95%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.