BioMiner: A Multi-modal System for Automated Mining of Protein-Ligand Bioactivity Data from Literature
Yan, J.; Zhu, J.; Yang, Y.; Liu, Q.; Zhang, K.; Zhang, Z.; Liu, X.; Zhang, B.; Gao, K.; Xiao, J.; Chen, E.
Show abstract
Protein-ligand bioactivity data published in literature are essential for drug discovery, yet manual curation struggles to keep pace with rapidly growing literature. Automated bioactivity extraction remains challenging because it requires not only interpreting biochemical semantics distributed across text, tables, and figures, but also reconstructing chemically exact ligand structures (e.g. Markush structures). To address this bottleneck, we introduce BO_SCPLOWIOC_SCPLOWMO_SCPLOWINERC_SCPLOW, a multi-modal extraction framework that explicitly separates bioactivity semantic interpretation from ligand structure construction. Within BO_SCPLOWIOC_SCPLOWMO_SCPLOWINERC_SCPLOW, bioactivity semantics are inferred through direct reasoning, while chemical structures are resolved via a chemical-structure-grounded visual semantic reasoning (CSG-VSR) paradigm, in which multi-modal large language models operate on chemically grounded visual representations to infer inter-structure relationships, and exact molecular construction is delegated to domain chemistry tools. For rigorous evaluation and method development, we further establish BO_SCPLOWIOC_SCPLOWVO_SCPLOWISTAC_SCPLOW, a comprehensive benchmark comprising 16,457 bioactivity entries curated from 500 publications. On BO_SCPLOWIOC_SCPLOWVO_SCPLOWISTAC_SCPLOW, BO_SCPLOWIOC_SCPLOWMO_SCPLOWINERC_SCPLOW validates its extraction ability and provides a quantitative baseline, achieving F1 scores of 0.33 for complex bioactivity triplets. BO_SCPLOWIOC_SCPLOWMO_SCPLOWINERC_SCPLOWs practical utility is demonstrated via three applications: (1) extracting 82,262 data from 11,683 papers to build a pre-training database that improves downstream models performance by 3.9%; (2) enabling a human-in-the-loop workflow that doubles the number of high-quality NLRP3 bioactivity data, helping 38.6% improvement over 28 QSAR models and identification of 16 hit candidates with novel scaffolds; and (3) accelerating the annotation of the protein-ligand structures in PoseBusters benchmark with reported bioactivity, achieving a 5-fold speed increase and 10% accuracy improvement over manual methods. BO_SCPLOWIOC_SCPLOWMO_SCPLOWINERC_SCPLOW and BO_SCPLOWIOC_SCPLOWVO_SCPLOWISTAC_SCPLOW provide a scalable extraction methodology and a rigorous benchmark, paving the way to unlock bioactivity data that previously required extensive human effort. All codes and data are available at GitHub.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data 96%
- Deep learning molecular interaction motifs from Receptor structure alone 95%
- Chemical Genomics Language Model toward Reliable and Explainable Compound-Protein Interaction Exploration 95%
Similar papers in this journal
- A Transferable Deep Learning Approach to Fast Screen Potent Antiviral Drugs against SARS-CoV-2 95%
- Enhanced compound-protein binding affinity prediction by representing protein multimodal information via a coevolutionary strategy 94%
- Revolutionizing GPCR-Ligand Predictions: DeepGPCR with experimental Validation for High-Precision Drug Discovery 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.