NeoGx: Machine-Recommended Rapid Genome Sequencing for Neonates
Antoniou, A. A.; McGinley, R.; Metzler, M.; Chaudhari, B. P.
Show abstract
BackgroundGenetic disease is common in the Level IV Neonatal Intensive Care Unit (NICU), but neonatology providers are not always able to identify the need for genetic evaluation. We trained a machine learning (ML) algorithm to predict the need for genetic testing within the first 18 months of life using health record phenotypes. MethodsFor a decade of NICU patients, we extracted Human Phenotype Ontology (HPO) terms from clinical text with Natural Language Processing tools. Considering multiple feature sets, classifier architectures, and hyperparameters, we selected a classifier and made predictions on a validation cohort of 2,241 Level IV NICU admits born 2020-2021. ResultsOur classifier had ROC AUC of 0.87 and PR AUC of 0.73 when making predictions during the first week in the Level IV NICU. We simulated testing policies under which subjects begin testing at the time of first ML prediction, estimating diagnostic odyssey length both with and without the additional benefit of pursuing rGS at this time. Just by using ML to accelerate initial genetic testing (without changing the tests ordered), the median time to first genetic test dropped from 10 days to 1 day, and the number of diagnostic odysseys resolved within 14 days of NICU admission increased by a factor of 1.8. By additionally requiring rGS at the time of positive ML prediction, the number of diagnostic odysseys resolved within 14 days was 3.8 times higher than the baseline. ConclusionsML predictions of genetic testing need, together with the application of the right rapid testing modality, can help providers accelerate genetics evaluation and bring about earlier and better outcomes for patients.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models 94%
- Reducing Sanger Confirmation Testing through False Positive Prediction Algorithms 93%
- One in seven pathogenic variants can be challenging to detect by NGS: An analysis of 450,000 patients with implications for clinical sensitivity and genetic test implementation 92%
Similar papers in this journal
Similar papers in this journal
- From Text to Translation: Using Language Models to Prioritize Variants for Clinical Review 93%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 92%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 91%
Similar papers in this journal
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 93%
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 93%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 92%
Similar papers in this journal
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 93%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 93%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.