Back

JIND-Multi: Leveraging Multiple Labeled Datasets for Automated Annotation of Single-Cell RNA and ATAC Data

Sancho Zamora, J.; Kanhirodan, A.; Garrote, X.; Gevaert, O.; Hernaez, M.; Serrano, G.; Ochoa, I.

2025-01-19 bioinformatics
10.1101/2025.01.15.633130 bioRxiv
Show abstract

BackgroundThe creation of single-cell atlases is essential for understanding cellular diversity and heterogeneity. However, assembling these atlases is challenging due to batch effects and the need for accurate cell annotation. Current methods for single-cell RNA and ATAC sequencing, while effective for integration, are not optimized for cell annotation. Additionally, many annotation tools rely on external databases or reference scRNA-Seq datasets, which may limit their adaptability to specific study needs, especially for rare cell-types or scATAC-Seq data. ResultsWe introduce JIND-Multi, an extended version of the JIND framework, designed to transfer cell-type labels across multiple annotated datasets. JIND-Multi significantly reduces the proportion of unclassified cells in single-cell RNA sequencing (scRNA-Seq) data while maintaining the accuracy and performance of the original JIND model. Furthermore, JIND-Multi demonstrates robust and precise annotation results in its inaugural application to scATAC-Seq data, proving its versatility and effectiveness across different single-cell sequencing technologies. ConclusionsJIND-Multi represents an improvement in cell annotation, reducing unassigned cells and offering a reliable solution for both scRNA-Seq and scATAC-Seq data. Its ability to handle multiple labeled datasets enhances the precision of annotations, making it a valuable tool for the single-cell research community. JIND-Multi is publicly available at:https://github.com/ML4BM-Lab/JIND-Multi.git.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.