Back

Analysis of a Large Patient-Level Dataset to Predict Outcome of Treatment for Drug-Resistant Tuberculosis

Wang, Q.; Gu, J.; Gabrielian, A.; Rosenfeld, G.; Quinones, M.; Hurt, D. E.; Rosenthal, A.

2022-09-17 health informatics
10.1101/2022.09.14.22279738 medRxiv
Show abstract

BACKGROUNDDrug-resistant (DR) tuberculosis treatment is challenging and frequently leads to poor outcomes. An international collaboration, the National Institute of Allergy and Infectious Diseases (NIAID) TB Portals develops, maintains, and supports a multi-national database of tuberculosis cases, with an emphasis on drug-resistant tuberculosis. Patient records include clinical, radiological, genomic, and socioeconomic features. Establishing factors associated with unsuccessful treatment may help optimize treatment for the most challenging infections. METHODSAssociation analysis and machine learning algorithms were applied to identify important factors associated with treatment outcome and predict the outcome for three patient cohorts, selected by drug resistance level representing 1575 patients in total. The predicted probabilities of poor treatment outcome from models were calibrated as a risk score ranging from 0 to 100 corresponding to confidence level of the model for treatment outcome. RESULTSThe features most associated with treatment success in all cohorts were body mass index (BMI), onset age, employment, education, smear-negative microscopy, and percent of abnormal volume in X-ray images, confirming previously reported findings, and identifying novel factors such as pathogen genomic markers. CONCLUSIONSThe identified features might help in establishing high-risk patients at the time of admission for tuberculosis treatment. This study integrates clinical, radiological, and pathogen genomics into a patient risk model, a way of determining risk through the application of machine learning on real-world data.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.