Back

Combining machine learning algorithms and single-cell data to study the pathogenesis of Alzheimer's disease

Xie, L. G.; Cui, W.; Lee, E.

2024-01-29 bioinformatics
10.1101/2024.01.26.577320 bioRxiv
Show abstract

Extracting valuable insights from high-throughput biological data of Alzheimers disease to enhance understanding of its pathogenesis is becoming increasingly important. We engaged in a comprehensive collection and assessment of Alzheimers microarray datasets GSE5281 and GSE122063 and single-cell data from GSE157827 from the NCBI GEO database. The datasets were selected based on stringent screening criteria: a P-value of less than 0.05 and an absolute log fold change (|logFC|) greater than 1. Our methodology involved utilizing machine learning algorithms, efficiently identified characteristic genes. This was followed by an in-depth immune cell infiltration analysis of these genes, gene set enrichment analysis (GSEA) to elucidate differential pathways, and exploration of regulatory networks. Subsequently, we applied the Connectivity Map (cMap) approach for drug prediction and undertook single-cell expression analysis. The outcomes revealed that the top four characteristic genes, selected based on their accuracy, exhibited a profound correlation with the Alzheimers disease (AD) group in terms of immune infiltration levels and pathways. These genes also showed significant associations with multiple AD-related genes, enhancing the potential pathogenic mechanisms through regulatory network analysis and single-cell expression profiling. Identified three subpopulations of astrocytes in late-stage of AD Prefrontal cortex dataset. Discovering dysregulation of the expression of the AD disease-related pathway maf/nrf2 in these cell subpopulations Ultimately, we identified a potential therapeutic drug score, offering promising avenues for future Alzheimers disease treatment strategies.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.