Back

TubercuProbe: A Cross-Attention Graph-Sequence Model for Cross-Species Chemoproteomic Discovery in Mycobacterium tuberculosis

Chalamalasetty, A.; Mishra, A. R.; Ferguson, F.; Sanchez-Lengeling, B.; Barraza-Chavez, J. M.; Jinich, A.

2026-01-05 biochemistry
10.64898/2026.01.04.697588 bioRxiv
Show abstract

Activity-based protein profiling (ABPP) and residue-specific chemoproteomics have transformed human chemical biology, yet applying these approaches to pathogens remains limited by biosafety constraints, low throughput, and the absence of reusable atlases. We present TubercuProbe, a cross-species machine learning framework that leverages large-scale human chemoproteomic knowledge to prioritize compound-protein interactions in Mycobacterium tuberculosis (Mtb) and other pathogens. Our model integrates a graph isomorphism network (GINE) for ligand encoding with frozen ESM-C (600M) protein embeddings via bidirectional cross-attention. Trained on >2M ChEMBL compound-protein pairs (predominantly human targets), TubercuProbe achieves R2=0.77 (MSE=0.45) for continuous affinity prediction and transfers effectively to binary cysteine reactivity prediction (CysDB AUPRC=0.63). Ablation studies reveal that pretrained features are highly transferable across freezing strategies ({Delta}AUPRC<0.03), suggesting the model captures fundamental protein-molecule interaction patterns. As a case study, we prioritize cysteine-reactive electrophiles and molecular glues for three Mtb virulence proteins (PtpB, SapM, Rv3671c), providing candidate probes for prospective ABPP validation. Orthogonal comparison with Boltz-2 structure predictions shows moderate correlation (Pearson r{approx}0.69). TubercuProbe provides a lightweight, sequence-driven first-pass ranker that enables pre-experimental prioritization--reducing the time, cost, and experimental burden of chemoproteomic discovery in biosafety-restricted systems. We discuss extensions toward multitask learning that jointly predicts non-covalent binding and covalent reactivity, recognizing that effective covalent probes must both reach their target and react once there.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.