EpiKG2DAG: a Framework for Automated DAG Construction from Biomedical Text
DU, J.; Deng, G.
Show abstract
While Directed Acyclic Graphs (DAGs) are essential for causal inference, their construction often relies on expert heuristics, which bypasses systematic evidence synthesis and creates a critical "evidence retrieval gap" in causal modeling. This study introduces EpiKG2DAG, a framework that supports evidence-anchored candidate DAG generation by transforming unstructured biomedical abstracts into structured epidemiological associations. We utilized DeepSeek-V3 to extract exposure-outcome association triplets from 189,266 abstracts and employed SapBERT for semantic normalization against UMLS concepts. The resulting Epidemiological Knowledge Graph (EpiKG) enables the automated identification of candidate confounders, mediators, and colliders based on graph-theoretic motifs and literature-derived evidence. A case study on COVID-19 and AKI demonstrates that the framework uncovers non-obvious confounders, such as air pollution, while ensuring evidence traceability. This work contributes to the field by mitigating the knowledge-acquisition bottleneck and providing a transparent, reproducible foundation for evidence-based causal modeling.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance 94%
- Causal feature selection using a knowledge graph combining structured knowledge from the biomedical literature and ontologies: a use case studying depression as a risk factor for Alzheimer's disease 93%
- Creating an Ignorance-Base: Exploring Known Unknowns in the Scientific Literature 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Trialstreamer: a living, automatically updated database of clinical trial reports 92%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 92%
- medExtractR: A medication extraction algorithm for electronic health records using the R programming language 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.