Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification
Reisle, C.; Grisdale, C. J.; Krysiak, K.; Danos, A. M.; Khanfar, M.; Pleasance, E.; Saliba, J.; Hanos, M.; Patel, N. V.; Jain, A.; McMichael, J. F.; Venigalla, A. C.; Griffith, M.; Griffith, O. L.; Jones, S. J. M.
Show abstract
Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Platform for Oncogenomic Reporting and Interpretation 94%
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 94%
- Deep representation learning for clustering longitudinal survival data from electronic health records 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 92%
- Deep integrative models for large-scale human genomics 92%
- Interpretable deep learning for chromatin-informed inference of transcriptional programs driven by somatic alterations across cancers 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.