Back

PhenoXtract: combining Large Language Model and Knowledge Graph embedding to extract phenotypes from clinical descriptions

Berardelli, S.; BRIERE, G.; Loire, B.; De Paoli, F.; Gazzo, A. M.; Limongelli, I.; Magni, P.; Zucca, S.; Baudot, A.

2026-06-26 genomics
10.64898/2026.06.22.733382 bioRxiv
Show abstract

Motivation: Standardized phenotypic descriptions are essential for accurate diagnosis, yet clinicians and researchers face challenges in manually extracting and mapping phenotypes from scientific literature or patient clinical records to the Human Phenotype Ontology. Recent advances in deep learning offer new opportunities for automation. We developed PhenoXtract, a novel phenotype extraction approach that combines Large Language Models and Knowledge Graph embedding. PhenoXtract is a multistep pipeline that takes clinical descriptions as input, extracts candidate phenotype entities using large language models, and maps them to terms from an enriched version of the Human Phenotype Ontology, processed as a knowledge graph. Results: Evaluation against expert-curated ground-truth datasets show a recall of 0.70 and precision of 0.85 for PhenoXtract, demonstrating concordance with manually extracted phenotypes, with a computation time of 10-20 seconds for each text analyzed. Moreover, PhenoXtract surpasses rule-based and deep learning-based state-of-the-art tools in two out of the three ground-truth datasets evaluated. These results suggest that hybrid approaches combining Large Language Models and Knowledge Graph embeddings represent a promising direction for automated clinical phenotyping at scale.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Genetics in Medicine
78 papers in training set
Top 0.1%
18.6%
2
Scientific Reports
3612 papers in training set
Top 15%
5.5%
3
Artificial Intelligence in Medicine
17 papers in training set
Top 0.1%
5.5%
4
Genome Medicine
183 papers in training set
Top 0.6%
5.5%
5
BioData Mining
22 papers in training set
Top 0.1%
4.1%
6
Bioinformatics
1204 papers in training set
Top 5%
3.4%
7
BMC Bioinformatics
457 papers in training set
Top 3%
3.3%
8
BMC Medical Genomics
50 papers in training set
Top 0.2%
3.2%
9
Journal of Biomedical Informatics
47 papers in training set
Top 0.5%
2.8%
50% of probability mass above
10
Nature Medicine
125 papers in training set
Top 0.8%
2.8%
11
Nature Communications
5641 papers in training set
Top 39%
2.5%
12
Bioinformatics Advances
203 papers in training set
Top 2%
2.4%
13
npj Digital Medicine
118 papers in training set
Top 2%
2.4%
14
Database
61 papers in training set
Top 0.3%
2.4%
15
eBioMedicine
183 papers in training set
Top 2%
2.0%
16
European Journal of Human Genetics
58 papers in training set
Top 0.6%
1.9%
17
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
18
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.7%
19
Human Mutation
34 papers in training set
Top 0.3%
1.7%
20
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.5%
21
PLOS ONE
5266 papers in training set
Top 53%
1.3%
22
Scientific Data
209 papers in training set
Top 2%
1.1%
23
GigaScience
212 papers in training set
Top 3%
1.1%
24
Genetic Epidemiology
55 papers in training set
Top 0.6%
1.0%
25
Frontiers in Genetics
230 papers in training set
Top 5%
1.0%
26
Genome Research
468 papers in training set
Top 6%
0.8%
27
Genome Biology
637 papers in training set
Top 8%
0.8%
28
BMC Genomics
406 papers in training set
Top 9%
0.6%
29
Communications Medicine
113 papers in training set
Top 6%
0.6%
30
The American Journal of Human Genetics
234 papers in training set
Top 3%
0.6%