Back

Natural language processing and expert follow-up establishes tachycardia association with CDKL5 deficiency disorder

Ivaniuk, A.; Bosselmann, C.; Zhang, X.; St John, M.; Taylor, S. C.; Krishnaswamy, G.; Milinovich, A.; Aziz, P. F.; Pestana-Knight, E.; Lal, D.

2023-06-29 genetic and genomic medicine
10.1101/2023.06.24.23291719 medRxiv
Show abstract

PurposeCDKL5 deficiency disorder (CDD) is a developmental and epileptic encephalopathy with multisystemic comorbidities. Cardiovascular involvement in CDD was shown in animal models but is yet poorly described in CDD cohorts. MethodsWe identified 38 individuals with genetically confirmed CDD through the Cleveland Clinic CDD specialty clinic and matched 190 individuals with non-genetic epilepsy to them as a comparison group. Natural language processing was applied to yield Human Phenotype Ontology (HPO) terms from medical records. We conducted HPO association testing and manual chart review to explore cardiovascular comorbidities associated with CDD. ResultsWe extracted 243,541 HPO terms from 30,512 medical encounters. Phenome-wide analysis confirmed well-established CDD phenotypes and identified association of tachycardia with CDD (OR 4.2, 95%CI 1.75-9.93, padj<0.001). We found a 99.6-fold enrichment of supraventricular tachycardia (SVT) in CDD encounter notes (padj < 0.001), which led to identification of two cases of fetal/neonatal onset SVT previously undescribed in CDD. Tachycardia in CDD individuals was associated with the presence of other autonomic symptoms (OR 5.63, 95%CI 1.08-40.3, p=0.038). ConclusionsCDD is associated with tachycardia, potentially including early-onset supraventricular tachycardia. Alongside prospective validation studies, semiautomated genotype-phenotype analysis with matched controls is a scalable, rapid, and efficient approach for validating known and identifying novel phenotype associations.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.