Human Phenotype Ontology (HPO) Mapper: Semantic Mapping of Clinical Findings to the Human Phenotype Ontology Using AI-Powered Embeddings and LLM-Based Quality Control
Kadhim, A. Z.; Green, Z.; Boags, A.; George, M.; Heinson, A.; Stammers, M.; Kipps, C. M.; Beattie, R. M.; Robinson, P. N.; Ashton, J. J.; Ennis, S.
Show abstract
O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/25342726v1_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@8edb94org.highwire.dtl.DTLVardef@f20105org.highwire.dtl.DTLVardef@21033corg.highwire.dtl.DTLVardef@15b865e_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOVISUAL ABSTRACT:C_FLOATNO C_FIG Structured phenotypic annotations linked to genetic data can drive diagnostic insight and therapeutic discovery in complex diseases. However, poor research access to the rich clinical data trapped in unstructured clinical records remains a significant barrier to phenotype-genotype integration. Here, we present Human Phenotype Ontology (HPO) Mapper, a scalable AI-assisted tool designed to ingest semantically structured clinical findings paired with anatomical region and accurately map them to HPO terms and associated genes. We applied HPO Mapper to two forms of standardised clinical input extracted from inflammatory bowel disease (IBD) patient records. The first data type consisted of paired clinical findings + anatomical regions derived from unstructured clinical reports and the second was standardised ICD-10 code-derived phenotypes. HPO Mapper achieved high semantic alignment and mapping accuracy for both data types (F1 = 0.85 {+/-} 0.05 and 0.84 {+/-} 0.03, respectively). Our publicly available tool enables real-time HPO mapping for clinical applications providing a foundation for scalable AI-driven phenotyping across diseases.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 94%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 92%
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 90%
Similar papers in this journal
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 93%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 94%
- A Platform for Oncogenomic Reporting and Interpretation 92%
- Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.