Back

Subphase-Labeled Mitotic Dataset for AI-powered Cell Division Analysis

Ivan, Z. Z.; Hirling, D.; Grexa, I.; Ammeling, J.; Micsik, T.; Dobra, K.; Kuthi, L.; Sukosd, F.; Aubreville, M.; Miczan, V.; Horvath, P.

2025-07-21 bioinformatics
10.1101/2025.07.17.665280 bioRxiv
Show abstract

Mitosis detection represents a critical task in the field of digital pathology, as determination of the mitotic index (MI) plays an important role in the tumor grading and prognostic assessment of patients. Manual determination of MI is a labor-intensive and time-consuming task for practitioners with rather high interobserver variability, thus, automation has become a priority. There has been substantial progress towards creating robust mitosis detection algorithms in recent years, primarily driven by the Mitosis Domain Generalization (MIDOG) challenges. In parallel, there has been growing interest in the molecular characterization of mitosis with the goal of achieving a more comprehensive understanding of its underlying mechanisms in a subphase-specific manner. Here, we introduce a new mitotic figure dataset annotated with subphase information based on the MIDOG++ dataset as well as a previously unrepresented tumor domain to enhance the diversity and applicability of the dataset. We envision a new perspective for domain generalization by improving the performance of models with subtyping mitotic cells into the 5 main stages of normal mitosis, complemented with an atypical mitotic class. We believe that our work broadens the horizon in digital pathology: subtyping information could provide useful help for mitosis detection, while also providing promising new directions in answering biological questions, such as molecular analysis of the subphases on a single cell level.

Published in Scientific Data (predicted rank #5) · training set

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.