Back

Deep learning enables cross-species annotation and attribution of ageing states in haematopoietic stem and immune cells

Zhao, S.; Zhang, B.; zhai, x.; yau, c.; Lio, P.; Nerlov, C.

2026-07-26 bioinformatics
10.64898/2026.07.24.730582 bioRxiv
Show abstract

Mouse single-cell ageing studies provide experimentally controlled age contrasts, but using mouse-labelled data to annotate human ageing states is limited by species, donor and assay effects in sparse transcriptomic and chromatin profiles. We developed a cross-species annotation workflow that treats mouse-to-human prediction as a target-validated domain-adaptation problem. The workflow uses orthologue-aligned features, a residual encoder, an age classifier and a species discriminator trained with two-phase adversarial optimisation, and couples prediction with stability-based gene attribution. In haematopoietic stem cells (HSCs), the model achieved held-out human AUROCs of 0.933 in scRNA-seq and 0.953 in scATAC-seq. In an independent CD8+ T-cell scRNA-seq setting, the held-out human AUROC was 0.941. Ablation analyses indicated that residual connections, ELU activation and two-phase training improved predictive performance and attribution stability. Consensus attributions from DeepLIFT, Integrated Gradients and saliency recovered conserved ageing-associated genes with greater cross-species overlap than differential expression alone. In a COVID-19 convalescent cohort, severe disease in younger adults was associated with a higher fraction of CD8+ cells classified as old-like by the pretrained model. These results support a reproducible framework for testing, interpreting and releasing cross-species single-cell ageing models, while highlighting the need for target-domain validation when mouse labels are transferred to human data.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.