Back

Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases

Nguyen, Q. M.; Chen, F.; Liu, C.; Campbell, I. M.; Zhang, G.; Wu, D.; Szigety, K. M.; Sheppard, S. E.; Ahimaz, P.; Ta, C. N.; Chung, W. K.; Weng, C.; Wang, K.

2026-03-02 health informatics
10.64898/2026.02.26.26347223 medRxiv
Show abstract

Clinical decision making often relies on expert judgment guided by established guidelines, which can be challenging to standardize and abstract to implement. For example, selecting between gene panels and whole exome/genome sequencing (WES/WGS) for rare disease diagnosis frequently requires interpretation of evidence-based recommendations from the American College of Medical Genetics and Genomics (ACMG) guideline. Traditional machine learning (ML) models predicting suitable genetic tests often face interpretability limitations. We hypothesize that large language models (LLMs) can be fine-tuned to "mimic" clinicians reasoning patterns by interpreting and applying clinical guidelines with chain-of-thought (CoT). We present RareDAI, an integrative approach that addresses this challenge by analyzing heterogeneous clinical data, including unstructured notes and structured Phecodes. Using seven domain-specific questions, we guide the Llama 3.1 and Qwen 3 models to generate structured CoT outputs. These outputs are refined via our proposed self-distillation fine-tuning (SDFT) approach, enabling the model to produce interpretable reasoning prior to recommendation. RareDAI outperforms traditional supervised fine-tuning and base LLMs (e.g., Llama 3.1, GPT-4) by up to 10-20% in all metrics (accuracy, precision, recall, and F1-score) on both in-house data and external data, effectively assisting clinicians in selecting between diagnostic modalities across healthcare systems.

Published in npj Digital Medicine (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.