ORBIT: Orthogonal Rotation for Biological Inter-species Transfer
Wissenberg, P.; Lee, J. M.; Mutwil, M.
Show abstract
MotivationCross-species gene embeddings are central to transferring functional annotations between species. A recent method demonstrated that species-specific STRING (PPI) network embeddings can be aligned across 1322 eukaryotes with autoencoders (FedCoder), but this approach is computationally expensive, depends on careful hyperparameter selection, leaves substantial room for improvement in cross-species retrieval quality, and has not been demonstrated on coexpression networks. ResultsWe introduce an alignment pipeline for cross-species coexpression network embeddings based on orthogonal Procrustes rotation. Species-specific Node2Vec embeddings of coexpression networks are aligned to a shared space using ortholog anchors from OrthoFinder, solved in closed form via Singular Value Decomposition (SVD). Applied to 153 plant species and 5.7 million genes, Procrustes alignment achieves four-fold higher cross-species Spearman correlation and consistently higher retrieval metrics than the SPACE autoencoder, while leaving within-species coexpression structure invariant (preservation ratio 1.000 against the unaligned baseline). The full alignment completes in under three minutes on a single CPU, and on downstream tasks, Procrustes embeddings improve within-species GO term prediction and outperform SPACE for cross-species GO transfer. Procrustes and sequence embeddings remain complementary for biological-process prediction, consistent with observations from SPACE. AvailabilityCode for producing the embeddings is made available at https://github.com/pwissenberg/orbit.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Neighborhood nonnegative matrix factorization identifies patterns and spatially-variable genes in large-scale spatial transcriptomics data 94%
- Cross-species imputation and comparison of single-cell transcriptomic profiles 93%
- A k-mer-based maximum likelihood method for estimating distances of reads to genomes enables genome-wide phylogenetic placement. 93%
Similar papers in this journal
- Benchmarking strategies for cross-species integration of single-cell RNA sequencing data 95%
- Deep generative model embedding of single-cell RNA-Seq profiles on hyperspheres and hyperbolic spaces 94%
- Constructing Ensemble Gene Functional Networks Capturing Tissue/condition-specific Co-expression from Unlabled Transcriptomic Data with TEA-GCN 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.