megaMine: a scalable, rule-based framework for mining gene-cancer-drug evidence from biomedical literature
JUNAID, M.; Prazanowska, K. H.; Jeong, H.-E.; Ryu, Y.; Choi, J.; An, J.-Y.; Lim, S. B.
Show abstract
The rapid expansion of the oncology literature has outpaced manual curation of clinically relevant gene-cancer-drug associations and oncogenic driver evidence. Existing automated approaches often lack transparency or are difficult to scale across heterogeneous data sources. To address this gap, we developed megaMine, a transparent, rule-based, and context-aware literature-mining framework that integrates therapeutic and driver evidence from PubMed, PubTator, and Europe PMC by combining entity recognition, hierarchical heuristics, and contextual labeling. In therapy mode, megaMine was applied to approximately 100,000 oncology articles published between 2015 and 2025, yielding more than 23,000 structured sentence-level evidence records, with standardized annotations for drug response, resistance, and study context. Internal evaluation of context labels showed the strong separability between efficacy and non-efficacy evidence using ridge logistic regression (AUROC = 0.915; AUPRC = 0.941). Benchmarking against NCI/OncoKB-supported drug-cancer associations showed that curated clinical associations had higher megaMine composite evidence scores than unlabeled comparison pairs [median (IQR): 25.6 (9.07-72.5) vs. 3.61 (1.69-8.69); Wilcoxon rank-sum test, P < 2.2 x 10-16]. In driver mode, megaMine retrieved mutation- and biomarker-related evidence from an ERBB-focused gastric cancer query, generating 750 evidence rows from 200 PMIDs. These results demonstrate that deterministic and interpretable approaches can support scalable evidence extraction for downstream applications such as knowledge graph construction and literature-based evidence synthesis.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Systematic identification of pan-cancer single-gene expression biomarkers in drug high-throughput screens 94%
- Prognostic pan-cancer and single-cancer models: A large-scale analysis using a real-world clinico-genomic database 94%
- GOcats: A tool for categorizing Gene Ontology into subgraphs of user-defined concepts 91%
Similar papers in this journal
- NCT Precision Oncology Thesaurus Drugs – a Curated Database for Drugs, Drug Classes, and Drug Targets in Precision Cancer Medicine 94%
- DeepPhe-CR: Natural Language Processing Software Services for Cancer Registrar Case Abstraction 92%
- Simple Linear Cancer Risk Prediction Models with Novel Features Outperform Complex Approaches 91%
Similar papers in this journal
- Trialstreamer: a living, automatically updated database of clinical trial reports 95%
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 94%
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 92%
Similar papers in this journal
- preon: Fast and accurate entity normalization for drug names and cancer types in precision oncology 94%
- Using Cancer Profiles to Identify Synthetic Lethal Therapeutic Targets and Predictive Biomarkers in Cancer Gene Dependency Data 93%
- AI-HOPE: An AI-Driven conversational agent for enhanced clinical and genomic data integration in precision medicine research 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.