Lense: Optimizing data preprocessing in single-cell omics using LLMs
Liu, J.; Ji, Z.
Show abstract
Data preprocessing is critical for single-cell omics analyses, but default pipelines often underperform on diverse datasets, especially from emerging platforms like spatial transcriptomics. We introduce Lense, a language-model-guided method that automatically selects optimal preprocessing by comparing plots that visualize low-dimensional representations across pipeline variants. Integrated with Seurat, Lense streamlines analysis and improves preprocessing robustness without requiring manual tuning. Biographical NoteJingyun Liu is a Masters student in the Department of Biostatistics and Bioinformatics at Duke University. Dr. Zhicheng Ji is a tenure-track Assistant Professor in the Department of Biostatistics and Bioinformatics at Duke University. His research focuses on artificial intelligence and statistical modeling for single-cell genomics, spatial genomics, and biomedical imaging.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ILoReg enables high-resolution cell population identification from single-cell RNA-seq data 96%
- Embeddings of genomic region sets capture rich biological associations in lower dimensions 96%
- Differential Expression Gene Explorer (DrEdGE): A tool for generating interactive online data visualizations for exploration of quantitative transcript abundance datasets 95%
Similar papers in this journal
Similar papers in this journal
- Single-cell Multi-omics Integration for Unpaired Data by a Siamese Network with Graph-based Contrastive Loss 96%
- Decoding Single-Cell Multiomics: scMaui - A Deep Learning Framework for Uncovering Cellular Heterogeneity in Presence of Batch Effects and Missing Data 96%
- CoSTA: Unsupervised Convolutional Neural Network Learning for Spatial Transcriptomics Analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.