Back

Evidence Aggregator: AI reasoning applied to rare disease diagnostics

Twede, H.; Conard, A. M.; Pais, L.; Bryen, S.; O'Heir, E.; Smith, G.; Paulsen, R.; Austin-Tse, C. A.; Bloemendal, A.; Simons, C.; Saponas, S.; Wander, J. D.; MacArthur, D. G.; Rehm, H. L.

2025-03-13 genomics
10.1101/2025.03.10.642480 bioRxiv
Show abstract

Variant assessment of rare disease diagnostics depends on using domain knowledge in the time- consuming process of retrieving, reviewing, and synthesizing clinical and technical information. To address these challenges, we developed the Evidence Aggregator (EvAgg), an open-source, generative-AI-based tool designed for rare disease diagnosis that systematically extracts relevant information from the scientific literature for any human gene. EvAgg provides a thorough and current summary of observed genetic variants and their associated clinical features, enabling rapid synthesis of evidence concerning gene-disease relationships. We constructed an expert-curated dataset and evaluated EvAggs performance. EvAgg achieves 92% recall in identifying relevant papers, 96% recall in detecting instances of genetic variation within those papers, and [~]80% accuracy in extracting individual case and variant-level content (e.g. zygosity, inheritance, variant type, and phenotype). Further, EvAgg complemented the process of manual literature review by identifying substantial additional relevant information. When tested with analysts in rare disease case analysis, EvAgg reduced review time by 34% (p-value < 0.002) and increased the number of papers, variants, and cases evaluated per unit time. These savings have the potential to reduce diagnostic latency and increase solve rates for challenging rare disease cases.

Published in Genetics in Medicine (predicted rank #1) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.