TubercuProbe: A Cross-Attention Graph-Sequence Model for Cross-Species Chemoproteomic Discovery in Mycobacterium tuberculosis
Chalamalasetty, A.; Mishra, A. R.; Ferguson, F.; Sanchez-Lengeling, B.; Barraza-Chavez, J. M.; Jinich, A.
Show abstract
Activity-based protein profiling (ABPP) and residue-specific chemoproteomics have transformed human chemical biology, yet applying these approaches to pathogens remains limited by biosafety constraints, low throughput, and the absence of reusable atlases. We present TubercuProbe, a cross-species machine learning framework that leverages large-scale human chemoproteomic knowledge to prioritize compound-protein interactions in Mycobacterium tuberculosis (Mtb) and other pathogens. Our model integrates a graph isomorphism network (GINE) for ligand encoding with frozen ESM-C (600M) protein embeddings via bidirectional cross-attention. Trained on >2M ChEMBL compound-protein pairs (predominantly human targets), TubercuProbe achieves R2=0.77 (MSE=0.45) for continuous affinity prediction and transfers effectively to binary cysteine reactivity prediction (CysDB AUPRC=0.63). Ablation studies reveal that pretrained features are highly transferable across freezing strategies ({Delta}AUPRC<0.03), suggesting the model captures fundamental protein-molecule interaction patterns. As a case study, we prioritize cysteine-reactive electrophiles and molecular glues for three Mtb virulence proteins (PtpB, SapM, Rv3671c), providing candidate probes for prospective ABPP validation. Orthogonal comparison with Boltz-2 structure predictions shows moderate correlation (Pearson r{approx}0.69). TubercuProbe provides a lightweight, sequence-driven first-pass ranker that enables pre-experimental prioritization--reducing the time, cost, and experimental burden of chemoproteomic discovery in biosafety-restricted systems. We discuss extensions toward multitask learning that jointly predicts non-covalent binding and covalent reactivity, recognizing that effective covalent probes must both reach their target and react once there.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adding stochastic negative examples into machine learning improves molecular bioactivity prediction 96%
- BOLD-GPCRs: A Transformer-Powered App for Predicting Ligand Bioactivity and Mutational Effects Across Class A GPCRs 95%
- Graph Attention Site Prediction (GrASP): Identifying Druggable Binding Sites Using Graph Neural Networks with Attention 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- DeepGraphMol, a multi-objective, computational strategy for generating molecules with desirable properties: a graph convolution and reinforcement learning approach 95%
- DrugDiff - small molecule diffusion model with flexible guidance towards molecular properties 95%
- All-Atom Protein Sequence Design using Discrete Diffusion Models 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.