InferPloidy: A fast ploidy inference tool accurately classifies cells with abnormal CNVs in large single-cell RNA-seq datasets
Sung, W.; Chae, J.; Moon, J.; Yoon, S.
Show abstract
BackgroundAccurate inference of copy number variation (CNV) and ploidy from single-cell RNA-seq data is essential for resolving tumor heterogeneity and identifying malignant cells, yet existing tools such as CopyKat and SCEVAN are limited by long runtimes and reduced accuracy in large or heterogeneous datasets. ResultsHere, we present InferPloidy, a high-speed and robust ploidy inference method built on InferCNV that combines graph-based cell-grouping with iterative Gaussian mixture modeling. Across multiple cancer types--breast cancer, non-small cell lung cancer, pancreatic ductal adenocarcinoma, and colorectal cancer--InferPloidy achieved up to two orders of magnitude faster runtimes than existing tools, while maintaining superior classification accuracy. This accurate separation of aneuploid tumor cells enabled the discovery of subtype-specific therapeutic targets, including ERBB2, ESR1, EGFR, and MET, as well as recurrent surfaceome markers such as CD82, F11R, SLC2A1, TM9SF2, CXADR, and PLPP2, several of which have preclinical or clinical relevance. ConclusionThese results establish InferPloidy as a scalable platform for CNV-guided tumor cell identification and surfaceome-based biomarker discovery, offering broad utility for precision oncology and translational research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- labelSeg: segment annotation for tumor copy number alteration profiles 96%
- Benchmarking copy number aberrations inference tools using single-cell multi-omics datasets 96%
- Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing 95%
Similar papers in this journal
- Identification and Utilization of Copy Number Information for Correcting Hi-C Contact Map of Cancer Cell Line 96%
- Probabilistic modeling methods for cell-free DNA methylation based cancer classification 95%
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 95%
Similar papers in this journal
Similar papers in this journal
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 95%
- Assessing the Performance of Methods for Cell Clustering from Single-cell DNA Sequencing Data 95%
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.