Identification of immune responses predictive of clinical protection against Malaria using novel statistical pipelines
Fonseca, A.; Biecek, P.; Cordeiro, C.; Sepulveda, N.
Show abstract
BackgroundNowadays, the chance of discovering the best antibody candidates for explaining naturally acquired protection to malaria and detecting exposure to malaria parasites has notably increased due to publicly available multi-sera data. The analysis of these data is typically divided into a feature selection phase followed by a predictive one where several models are constructed for the outcome of interest. A key question in the analysis is to determine which and how each feature should be included in the predictive stage. ResultsTo answer this question, we developed three approaches for classifying malaria protected and susceptible groups: (i) a basic and simple approach based on selecting antibodies via the nonparametric Mann-Whitney test; (ii) a dichotomization approach where each antibody was selected according to the optimal cut-off via maximization of the {chi}2 statistic for two-way tables; (iii) a hybrid parametric/non-parametric approach that integrates Box-Cox transformation followed by a t-test, together with the use of finite mixture models and the Mann-Whitney test as a last resort. We illustrated the application of these three approaches with published serological data for predicting clinical malaria in 121 Kenyan children. The predictive analysis was based on a Super-Learner where predictions from multiple classifiers were pooled together. Our results led to almost similar areas under the Receiver Operating Characteristic curves of 0.72 (95% CI = [0.61, 0.82]), 0.80 (95% CI = [0.71, 0.90]), 0.79 (95% CI = [0.7, 0.88]) for the simple, dichotomization and hybrid approaches, respectively. ConclusionsThe three feature selection strategies provided a better predictive performance of the outcome when compared to the previous results solely relying on Random Forests alone (AUC=0.68). Given the similar predictive performance, we recommended the three strategies should be used in conjunction in the same data set and selected according to their complexity.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of novel genetic variants in the malaria vaccine candidate PfRh5: structure-guided insights into potential function 93%
- The erythrocyte membrane properties of beta thalassaemia heterozygotes and their consequences for Plasmodium falciparum invasion 93%
- Plasmodium falciparum genomic surveillance reveals spatial and temporal trends, association of genetic and physical distance, and household clustering 93%
Similar papers in this journal
- Impact of a rapid decline in malaria transmission on antimalarial IgG subclasses and avidity 93%
- Identification of conserved cross-species B-cell linear epitopes in human malaria: A subtractive proteomics and immuno-informatics approach targeting merozoite stage proteins 91%
- Immune-Based Prediction of COVID-19 Severity and Chronicity Decoded Using Machine Learning 91%
Similar papers in this journal
- Evaluation of antibody serology to determine current helminth and Plasmodium falciparum infections in a co-endemic area in Southern Mozambique 93%
- Clinical performance validation of the STANDARD G6PD Test: A multi-country pooled analysis 92%
- Comparative clinical transcriptome of pir genes in severe Plasmodium vivax malaria 91%
Similar papers in this journal
Similar papers in this journal
- Predictive Immunoinformatics Reveal Promising Safety and Anti-Onchocerciasis Protective Immune Response Profiles to Vaccine Candidates (Ov-RAL-2 and Ov-103) in Anticipation of Phase I Clinical Trials 92%
- Plasmodium falciparum serology: A comparison of two protein production methods for analysis of antibody responses by protein microarray 92%
- A novel assessment method for COVID-19 humoral immunity duration using serial measurements in naturally infected and vaccinated subjects 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.