Generalizable Long COVID Subtypes: Findings from the NIH N3C and RECOVER Programs
Reese, J.; Blau, H.; Bergquist, T.; Loomba, J. J.; Callahan, T.; Laraway, B.; Antonescu, C.; Casiraghi, E.; Coleman, B.; Gargano, M.; Wilkins, K.; Cappelletti, L.; Fontana, T.; Ammar, N.; Antony, B.; Murali, T. M.; Karlebach, G.; McMurry, J. A.; Williams, A.; Moffitt, R.; Banerjee, J.; Solomonides, A. E.; Davis, H.; Kostka, K.; Valentini, G.; Sahner, D.; Chute, C. G.; Madlock-Brown, C.; Haendel, M. A.; Robinson, P. N.
Show abstract
Accurate stratification of patients with post-acute sequelae of SARS-CoV-2 infection (PASC, or long COVID) would allow precision clinical management strategies. However, the natural history of long COVID is incompletely understood and characterized by an extremely wide range of manifestations that are difficult to analyze computationally. In addition, the generalizability of machine learning classification of COVID-19 clinical outcomes has rarely been tested. We present a method for computationally modeling PASC phenotype data based on electronic healthcare records (EHRs) and for assessing pairwise phenotypic similarity between patients using semantic similarity. Our approach defines a nonlinear similarity function that maps from a feature space of phenotypic abnormalities to a matrix of pairwise patient similarity that can be clustered using unsupervised machine learning procedures. Using k-means clustering of this similarity matrix, we found six distinct clusters of PASC patients, each with distinct profiles of phenotypic abnormalities. There was a significant association of cluster membership with a range of pre-existing conditions and with measures of severity during acute COVID-19. Two of the clusters were associated with severe manifestations and displayed increased mortality. We assigned new patients from other healthcare centers to one of the six clusters on the basis of maximum semantic similarity to the original patients. We show that the identified clusters were generalizable across different hospital systems and that the increased mortality rate was consistently observed in two of the clusters. Semantic phenotypic clustering can provide a foundation for assigning patients to stratified subgroups for natural history or therapy studies on PASC.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 94%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 93%
- Long non-coding RNAs (lncRNAs) NEAT1 and MALAT1 are differentially expressed in severe COVID-19 patients: An integrated single cell analysis 93%
Similar papers in this journal
- Who is pregnant? defining real-world data-based pregnancy episodes in the National COVID Cohort Collaborative (N3C) 91%
- The Pandemic Response Commons 91%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 91%
Similar papers in this journal
Similar papers in this journal
- MOATAI-VIR - an AI algorithm that predicts severe adverse events and molecular features for COVID-19’s complications 95%
- Predicting bloodstream infection outcome using machine learning 92%
- Genome-wide investigation of gene-cancer associations for the prediction of novel therapeutic targets in oncology 92%
Similar papers in this journal
- Systematic analysis of genetic and phenotypic characteristics reveals antisense oligonucleotide therapy potential for one-third of neurodevelopmental disorders 91%
- Beyond gene-disease validity: capturing structured data on inheritance, allelic-requirement, disease-relevant variant classes, and disease mechanism for inherited cardiac conditions 90%
- CACTUS: integrating clonal architecture with genomic clustering and transcriptome profiling of single tumor cells 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.