GFETM: Genome Foundation-based Embedded Topic Model for scATAC-seq Modeling
Fan, Y.; Li, Y.; Ding, J.; Li, Y.
Show abstract
Single-cell Assay for Transposase-Accessible Chromatin with sequencing (scATAC-seq) has emerged as a powerful technique for investigating open chromatin landscapes at single-cell resolution. However, analyzing scATAC-seq data remain challenging due to its sparsity and noise. Genome Foundation Models (GFMs), pre-trained on massive DNA sequences, have proven effective at genome analysis. Given that open chromatin regions (OCRs) harbour salient sequence features, we hypothesize that leveraging GFMs sequence embeddings can improve the accuracy and generalizability of scATAC-seq modeling. Here, we introduce the Genome Foundation Embedded Topic Model (GFETM), an interpretable deep learning framework that combines GFMs with the Embedded Topic Model (ETM) for scATAC-seq data analysis. By integrating the DNA sequence embeddings extracted by a GFM from OCRs, GFETM demonstrates superior accuracy and generalizability and captures cell-state specific TF activity both with zero-shot inference and attention mechanism analysis. Finally, the topic mixtures inferred by GFETM reveal biologically meaningful epigenomic signatures of kidney diabetes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 98%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 97%
- OmicVerse: A single pipeline for exploring the entire transcriptome universe 97%
Similar papers in this journal
Similar papers in this journal
- scAlign: a tool for alignment, integration and rare cell identification from scRNA-seq data 97%
- CMOT: Cross Modality Optimal Transport for multimodal inference 97%
- Neighborhood nonnegative matrix factorization identifies patterns and spatially-variable genes in large-scale spatial transcriptomics data 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.