RVQ-Alpha: Bridging Single-Cell Transcriptomics and Large Language Models via Discrete Tokenization and Verifiable Reinforcement Learning
Li, G.; You, Y.; Fu, Y.; Zhou, W.; Tang, F.; Kong, J.; Tian, L.
Show abstract
Integrating continuous single-cell gene expression profiles with discrete-token large language models (LLMs) remains an open challenge: text-based methods are token-inefficient and discard quantitative precision, continuous embeddings preclude autoregressive generation, and single-codebook vector quantization cannot stratify biological scales. We present RVQ-Alpha, an end-to-end framework that bridges this modality gap through three contributions: O_LIa Residual Vector Quantization (RVQ) tokenizer that compresses each cell into a fixed 10-token sequence via eight residual codebooks embedded directly in the LLM vocabulary, requiring 3.4x fewer tokens than prior discrete methods and enabling bidirectional unification of cell interpretation and generation; C_LIO_LIscCoT-Synth, a teacher-student engine that grounds newly added biological tokens through evidence-before-conclusion reasoning, using the language modeling objective as the cross-modal alignment signal without a separate projection network; C_LI and (iii) a Fact-Aware RLVR system combining an ontology-grounded answer judge with saliency-weighted verification of biological claims against actual expression data, under dynamic gating that conditions hallucination suppression on task competence. Built on Qwen3-4B and trained via continued pretraining, supervised fine-tuning, and reinforcement learning, RVQ-Alpha substantially improves out-of-distribution generalization and rare-cell recognition across eight held-out datasets; ablations confirm that evidence-first grounding reduces hallucination more than fivefold.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.