Hierarchical cross-entropy loss improves atlas-scale single-cell annotation models
Cultrera di Montesano, S.; D'Ascenzo, D.; Raghavan, S.; Amini, A. P.; Winter, P.; Crawford, L.
Show abstract
Accurately annotating cell types is essential for extracting biological insight from single-cell RNA-seq data. Although cell types are naturally organized into hierarchical ontologies, most computational models do not explicitly incorporate this structure into their training objectives. We introduce a hierarchical cross-entropy loss that aligns model objectives with biological structure. Applied to architectures ranging from linear models to transformers, this simple modification significantly improves out-of-distribution performance (12-15%) without added computational cost.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 95%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 94%
- An efficient not-only-linear correlation coefficient based on machine learning 93%
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 95%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 95%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.