Back

Construction of a Standardized Time-Lapse Imaging Database and a Gradient Boosting Ensemble Framework for Integrating Zygote Morphokinetic Parameters with Conventional Embryo Assessment

ZHAO, M.; LIU, J.; HAN, D.; ZHANG, C.; ZHOU, Y.; CHEN, S.; LIU, C.

2026-08-24 obstetrics and gynecology
10.64898/2026.08.20.26359523 medRxiv
Show abstract

In vitro fertilization (IVF) laboratories equipped with timelapse incubators generate vast quantities of sequential embryo images, yet the absence of standardized, annotated databases impedes the development of reproducible computational tools for embryo assessment. Here we describe the construction of a standardized time-lapse imaging database comprising 631 normally fertilized zygotes from 218 treatment cycles, integrating timelapse image sequences, patient clinical records, and embryo developmental outcomes. We further present a gradient boosting decision tree (GBDT) ensemble framework that integrates zygote morphokinetic parameters-continuous time-series features extracted via a validated CNN-based segmentation algorithm (US Patent US11210494B2)-with conventional embryo assessment grades (categorical features per the Istanbul consensus). The fusion framework employs equal-weight initialization followed by iterative residual-decreasing training to optimally combine heterogeneous feature types. Ablation analysis demonstrated that the integrated model achieved an AUC of 0.78, significantly outperforming morphokinetics-only (AUC 0.71) and conventional-only (AUC 0.65) models, confirming the complementary value of the two data modalities. The database and fusion framework provide a reproducible foundation for embryo development assessment and are generalizable to other multimodal data integration tasks in reproductive medicine.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.