Back

Machine learning approaches for classificationof Plasmodium falciparum life cycle stagesusing single-cell transcriptomes

Shukla, S.; Choudhuri, S.; Iragavarapu, G. P.; Ghosh, B.

2022-07-24 bioinformatics
10.1101/2022.06.22.497155 bioRxiv
Show abstract

Malaria, spread by the female Anopheles mosquito, is a highly fatal disease widespread in many parts of the world, causing 0.4 million deaths globally. Vital gene expressions form the basis in the detection of malaria infection levels. Quantification of malaria parasite infected RBCs and classification of its life cycle stages are done at macroscopic level by experts, for making informed decisions. Off late multiple computational approaches have been proposed to circumvent the problem of dimensionality leading to accurate predicted results. In this work a dimensionality reduction technique based on Genetic Algorithm (GA) is applied on P. falciparum single-cell transcriptomics to arrive at an optimized subset of features from the larger dataset. Features are chosen based on their class variants considering increased efficiency and accuracy, to separately transform the selected elements into a lower dimension. For the classification of the life cycle of malaria parasite based on single cell transcriptome data, a three-pronged approach employing the multiclass Support Vector Machine (SVM), Logistic Regression (LR) and Random Forest (RF) techniques is used. Distribution of cells was visualised and mapped using the R-based Seurat package. Further, we constructed protein interaction networks of the genes identified by the feature selection method and elucidated the role of the proteins in progression of the parasite through its life cycle. Our approach presents a novel protocol to implement ML techniques on scRNA seq datasets and subsequently harnessing the extracted information for biomarker/drug target detection.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.