Back

A reproducibility-audit framework for generalizable versus dataset-specific molecular transition boundaries in Alzheimer's disease

Kim, Y.; Heo, W.; Park, S. J.; Kim, Y.; Cho, Y. E.

2026-09-01 neuroscience
10.64898/2026.08.24.746808 bioRxiv
Show abstract

Molecular staging of Alzheimer's disease (AD) increasingly defines transition boundaries along single-cell pseudo-progression trajectories, yet whether such boundaries reproduce across brain regions, cohorts and molecular modalities is rarely tested. We present a permutation-controlled audit that combines nine boundary-detection algorithms with a fixed marker panel and four orthogonal reproducibility axes-algorithmic consensus, region, cohort and modality. On synthetic data with planted ground-truth boundaries the audit reaches 100% sensitivity and 94% specificity, rejecting four distinct artefact classes each by a different axis. Applied to the Seattle Alzheimer's Disease Brain Cell Atlas middle temporal gyrus, it localizes a transition that is robust across algorithms and recovered in most cell types but does not generalize: its leading marker is attenuated or absent in prefrontal cortex, entorhinal cortex and cerebrospinal fluid, and an apparent cross-region conservation of glial metabolic genes proves to be a global-expression offset rather than a shared program. The same audit nonetheless certifies an externally validated marker (astrocytic PTGDS) as reproducible across regions and modalities, showing that it separates generalizable anchors from dataset-specific ones rather than rejecting all signals. We provide this four-axis audit as a transferable, code-available standard to apply before a trajectory boundary is read as a biological stage, in AD and other progressive proteinopathies.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.