Back

Validation of Real-World Case Definitions for COVID-19 Diagnosis and Severe COVID-19 Illness Among Patients Infected with SARS-CoV-2: Translation of Clinical Trial Definitions to Real-World Settings

Duh, M. S.; Nguyen, C.; Rubino, H.; Herrick, C.; Chang, R.; DerSarkissian, M.; Hsieh, Y. G.; Banatwala, A.; Yu, L. H.; Belsky, G.; Murphy, M. E.; Boyle-Kelly, J.; Cagan, A.; Stangle, B. E.; Cremieux, P. Y.; Kolitsopoulos, F.; Murphy, S. N.

2023-09-12 infectious diseases
10.1101/2023.09.12.23295441 medRxiv
Show abstract

PurposeThis study assessed the performance of International Classification of Diseases 10th Revision, Clinical Modification (ICD-10-CM) coronavirus disease 2019 (COVID-19) diagnostic code U07.1 against polymerase chain reaction (PCR) test results (Objective 1), and electronic medical record (EMR)-based codified algorithm for severe COVID-19 illness based on endpoints used in the Pfizer-BioNTech COVID-19 vaccine trial against chart review (Objective 2). MethodsThis retrospective, longitudinal cohort study used EMR data from the Mass General Brigham COVID-19 Data Mart (3/1/2020-11/19/2020) for adult patients with [≥]1 PCR test, antigen test, or code U07.1 (Objective 1) and adult patients with a positive PCR test hospitalized with COVID-19 (Objective 2). ResultsAmong 354,124 patients in Objective 1, 96% had [≥]1 PCR test (including 6% with [≥]1 positive PCR test; 11% with [≥]1 code U07.1). Code U07.1 had low sensitivity (54%) and positive predictive value (PPV; 63%) but high specificity (97%) against the PCR test. Among 300 patients hospitalized for COVID-19 randomly sampled for chart review in Objective 2, the EMR-based case definition for severe COVID-19 illness had high PPV (>95%), showing better performance than severe/critical COVID-19 endpoints defined by the World Health Organization (PPV: 79%). ConclusionsCOVID-19 diagnosis based on ICD-10-CM code U07.1 had inadequate sensitivity and requires confirmation by PCR testing. The EMR-based case definition showed high PPV and can be used to identify cases of severe COVID-19 illness in real-world datasets. These findings highlight the importance of validating outcomes in real-world data, and can guide researchers analyzing COVID-19 data when PCR tests are not readily available. KEY POINTSO_LIThis study evaluated the performance of International Classification of Diseases 10th Revision, Clinical Modification (ICD-10-CM) codes and an electronic medical record (EMR)-based algorithm for identifying coronavirus disease 2019 (COVID-19) diagnosis and severe COVID-19 illness in real-world data. C_LIO_LIICD-10-CM code U07.1 for COVID-19 had low sensitivity and positive predictive value (PPV) against PCR tests. C_LIO_LIThe EMR-based algorithm for severe COVID-19 illness developed from the Pfizer- BioNTech COVID-19 vaccine trial had high PPV against chart review, and may be used to identify severe cases in real-world data. C_LIO_LIThese results highlight the importance of validating outcomes when conducting analyses of real-world datasets. C_LI PLAIN LANGUAGE SUMMARYAs polymerase chain reaction (PCR) tests for coronavirus disease 2019 (COVID-19) diagnosis are becoming less frequently used and there is no standard definition of severe COVID-19 illness, it is important to have a way of correctly identifying COVID-19 diagnosis or severe COVID-19 illness in real-world data (e.g., electronic medical records [EMRs]). This study examined: 1) how a diagnosis code for COVID-19 used in EMRs (i.e., U07.1) compares to PCR test results in terms of accurately identifying patients with COVID-19; and 2) whether a definition for severe COVID-19 illness developed based on the Pfizer-BioNTech COVID-19 vaccine trial and a definition used by the World Health Organization [WHO]) can be used to accurately identify patients with severe COVID-19 illness in EMRs. The results showed that code U07.1 was not very accurate in identifying patients with COVID-19. On the other hand, the developed definition for severe COVID-19 illness was more accurate than the WHO definition and was able to identify most patients with severe COVID-19 illness in real-world data.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.