CellTosg2Sequence: A Unified Text-Omics-Signaling-Graph Large Language Model for Single-Cell Analysis
chen, w.; Ye, M.; Xu, T.; Huang, D.; Zhang, H.; Li, H.; Li, W.; Chen, Y.; Payne, P. R.; Li, F.
Show abstract
In single-cell (sc)-based scientific discovery, text-formatted biomedical prior knowledge and signaling graphs are essential for annotating and interpreting numeric sc-omics data and for generating novel testable hypotheses. A major limitation of existing single-cell large language models (scLLMs) is that they rely on numeric expression data with gene names as the only textual signal, while comprehensive biomedical priors -- cellular localization, gene function, disease associations, and signaling interaction patterns -- remain absent from the model input. We introduce CellTosg2Sequence, a textual-prior- and signaling-graph-augmented cell-omics-sentence language model. A lightweight heterogeneous graph encoder maps a curated 62,507-node biomedical knowledge graph (KG) into compact virtual tokens that are prepended to each cell sentence, allowing the language model to condition on biological structure with minimal sequence-length overhead. We train CellTosg2Sequence with a three-stage objective: Stage I anchors the KG channel under autoregressive language-model pretraining, leveraging Qwen2.5-32Bs own language reasoning for rapid KG alignment; Stage II aligns labels via supervised fine-tuning with KG-anchored InfoNCE; Stage III applies Group Relative Policy Optimization (GRPO) with an ontology-hierarchy reward, enabling free-generation cell-type prediction that generalizes beyond the closed training vocabulary. Across multiple benchmarks and ablation experiments, CellTosg2Sequence outperforms strong baselines. All results are achieved with lightweight LoRA training and a single unified checkpoint. Data ethicsThis work uses publicly available single-cell datasets from the Human Cell Atlas (https://www.humancellatlas.org) and the Tahoe-100M consortium. All HCA constituent studies were collected under appropriate donor consent and institutional oversight as described in their original publications; we perform computational re-analysis only and introduce no new human subjects data. HCA data access follows the HCA Data Portal terms of use. No new patient data are collected in this study; no additional IRB approval is required for this secondary computational analysis.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Sampling from Disentangled Representations of Single-Cell Data Using Generative Adversarial Networks 95%
- Enhancement of network architecture alignment in comparative single-cell studies 95%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.