PhenoID, a language model normalizer of physical examinations from genetics clinical notes
Weissenbacher, D.; Rawal, S.; Zhao, X.; Priestley, J. R. C.; Szigety, K. M.; Schmidt, S. F.; Higgins, M. J.; Magge, A.; O'Connor, K.; Gonzalez-Hernandez, G.; Campbell, I. M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSPhenotypes identified during dysmorphology physical examinations are critical to genetic diagnosis and nearly universally documented as free-text in the electronic health record (EHR). Variation in how phenotypes are recorded in free-text makes large-scale computational analysis extremely challenging. Existing natural language processing (NLP) approaches to address phenotype extraction are trained largely on the biomedical literature or on case vignettes rather than actual EHR data. MethodsWe implemented a tailored system at the Childrens Hospital of Philadelpia that allows clinicians to document dysmorphology physical exam findings. From the underlying data, we manually annotated a corpus of 3136 organ system observations using the Human Phenotype Ontology (HPO). We provide this corpus publicly. We trained a transformer based NLP system to identify HPO terms from exam observations. The pipeline includes an extractor, which identifies tokens in the sentence expected to contain an HPO term, and a normalizer, which uses those tokens together with the original observation to determine the specific term mentioned. FindingsWe find that our labeler and normalizer NLP pipeline, which we call PhenoID, achieves state-of-the-art performance for the dysmorphology physical exam phenotype extraction task. PhenoIDs performance on the test set was 0.717, compared to the nearest baseline system (Pheno-Tagger) performance of 0.633. An analysis of our systems normalization errors shows possible imperfections in the HPO terminology itself but also reveals a lack of semantic understanding by our transformer models. InterpretationTransformers-based NLP models are a promising approach to genetic phenotype extraction and, with recent development of larger pre-trained causal language models, may improve semantic understanding in the future. We believe our results also have direct applicability to more general extraction of medical signs and symptoms. FundingUS National Institutes of Health
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natural language inference for clinical registry curation 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Medication information extraction using local large language models 96%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 94%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.