Neural network-based identification of easily-obtainable demographic and clinical characteristics to identify people with tuberculosis
Parihar, D. S.; Jansen van Vüren, J. M.; Naidoo, D.; Klopper, M.; Cobelens, F.; Zajac, K.; Kolbe, L. M.; Ssengooba, W.; Joloba, M.; Theron, G.; Niesler, T.
Show abstract
We consider the application of machine learning to the classification of tuberculosis (TB) based on clinical and demographic data. Such data is routinely collected from people who present with cough at community-level care centres. Therefore, such automatic classification could identify people who require expensive but critical confirmatory testing, thereby offering simple and low-cost method of triage. Logistic regression, XGBoost, and convolutional neural network classifiers are evaluated using fully-nested cross validation, with and without feature selection. Although the application of CNNs to clinical and demographic data is unconventional, we show it to be effective. Experiments are carried out using two datasets: cough diagnostic algorithm for TB (CODA TB), n = 1140 and cough audio triage for TB (CAGE-TB), n = 463, for both datasets all participants self-presented to healthcare facilities with symptoms or risk factors suggestive of TB. Using the CNN, areas under the receiver operating characteristic (AUROC) of 80.48% and 83.06% are achieved for the two datasets respectively. Furthermore, performance is shown to improve both when the set of clinical features is extended, and when the number of people in the dataset increases. This holds promise of the development of an automated TB triage tool, implemented on a low-cost mobile device such as a smartphone, that is suitable for use at primary health-care facilities.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 93%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 92%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
Similar papers in this journal
- Multinomial modelling of TB/HIV co-infection yields a robust predictive signature and generates hypotheses about the HIV+TB+ disease state. 93%
- Cohort profile: St. Michael's Hospital Tuberculosis Database (SMH-TB), a retrospective cohort of electronic health record data and variables extracted using natural language processing 92%
- Deep learning models for COVID-19 chest x-ray classification: Preventing shortcut learning using feature disentanglement 92%
Similar papers in this journal
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 95%
- Data-driven identification of previously unrecognized communities with alarming levels of tuberculosis infection in the Democratic Republic of Congo 94%
- Development and Clinical Validation of Swaasa AI Platform for screening and prioritization of Pulmonary TB 94%
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 94%
- COVID-19 diagnosis prediction by symptoms of tested individuals: a machine learning approach 91%
- Passive Detection of COVID-19 with Wearable Sensors and Explainable Machine Learning Algorithms 91%
Similar papers in this journal
- Predicting the causative pathogen among children with pneumonia using a causal Bayesian network 92%
- Using genetic data to identify transmission risk factors: statistical assessment and application to tuberculosis transmission 92%
- Cluster detection with random neighbourhood covering: application to invasive Group A Streptococcal disease 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.