Back

A Blood-Based Transcriptomic Algorithm and Scoring System for Alzheimer's Disease Detection

Goyal, P.; Arya, N.

2025-10-22 health informatics
10.1101/2025.10.21.25338438 medRxiv
Show abstract

Alzheimers disease is a neurodegenerative disorder that affects more than 50 million people worldwide. Current diagnostic methods include cerebrospinal fluid testing and CT/MRI scans, which are either invasive, expensive, or non-specific. This study aimed to develop alternative diagnostic approaches for early detection by creating a transparent algorithm using blood-based transcriptomic biomarkers associated with Alzheimers. Towards this, a microarray dataset from NCBI GEO was obtained, which contained the expression levels of 14,113 genes in 180 subjects: 90 Alzheimers cases and 90 non-Alzheimers controls. The Mann-Whitney U Test and the Holm-Bonferroni p-value correction were applied for feature selection, yielding 8 statistically significant genes. Further, symbolic regression using the Quantum Lattice technique led to the generation of a 3-gene mathematical function that could be utilized for Alzheimers prediction. The three genes identified in the regression model were FBRSL1, TRIB2, and LY6G6D, the first two being novel discoveries. Thereafter, the results were converted to a 100-point scoring system intended for clinical use, and the diagnostic scores were validated using the testing dataset. The model assigned scores lower than 20 to non-Alzheimers controls with 92% accuracy and scores greater than 80 to Alzheimers patients with 75% accuracy. The scoring system is a successful proof of concept that can serve as a starting point for more accurate risk systems in the future, and the identified biomarkers can be further validated in other cohorts of Alzheimers patients.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.