PubMind: Literature-Based Genetic Variant Extraction and Functional Annotation Using Large Language Models
Wang, P.; Wang, K.
Show abstract
The rapid growth of biomedical literature has produced extensive functional knowledge on genetic variants, much of which remains buried in unstructured texts. Current databases such as ClinVar and the Human Gene Mutation Database (HGMD) attempt to catalog this knowledge but have significant limitations: ClinVar depends on voluntary submissions and covers only a fraction of published literature, while the academic version of HGMD is updated infrequently and provides limited functional annotation. To address these gaps, we developed PubMind, an AI-driven multi-layer framework that uses large language models (LLMs) to extract variant- function-disease associations and supporting evidence from text. PubMind integrates a fine-tuned BERT model for input triage with instruction-tuned GPT models for inferring disease associations and functional annotations. The system captures diverse variant types--including SNVs, CNVs, SVs, and gene fusions--and normalizes records to genome and transcriptome coordinates. Benchmarking demonstrates >90% accuracy in variant recognition and 99% precision in disease extraction. Application of PubMind on >41 million PubMed abstracts and >5 million open-access full-text articles produced PubMind-DB, a database containing [~]1.3 million unique variants with rich contextual annotations, accessible via a web interface and API. Only [~]10% of PubMinds variants overlapped with ClinVar entries, yet >80% showed concordant pathogenicity labels, including full agreement with ClinVars expert-reviewed variants. Case studies demonstrate PubMind-DBs ability to uncover supporting evidence for variant pathogenicity that might otherwise be missed by manual searches. Together, these findings establish PubMind as a scalable LLM-based framework that transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 97%
- Diverse ancestral representation improves genetic intolerance metrics 95%
- Saturation genome editing of DDX3X clarifies pathogenicity of germline and somatic variation 95%
Similar papers in this journal
- Genomics 2 Proteins portal: A resource and discovery tool for linking genetic screening outputs to protein sequences and structures 96%
- Morphological map of under- and over-expression of genes in human cells 94%
- Haplotype-aware variant calling enables high accuracy in nanopore long-reads using deep neural networks 94%
Similar papers in this journal
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 97%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 96%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 95%
Similar papers in this journal
Similar papers in this journal
- AutoPM3: Enhancing Variant Interpretation via LLM-driven PM3 Evidence Extraction from Scientific Literature 97%
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 95%
- Flashzoi: An enhanced Borzoi model for accelerated genomic analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.