Back

Comparing Expert and Computerised Pattern Identification in Antepartum Cardiotocography

Tolladay, J.; Albert, B.; Cooke, W.; Vatish, M.; Davis Jones, G.

2025-08-19 obstetrics and gynecology
10.1101/2025.08.15.25333416 medRxiv
Show abstract

ObjectiveTo assess the reliability of antepartum cardiotocography (CTG) pattern identi-fication by quantifying the level of agreement among expert clinicians and between clinicians and the Dawes-Redman (DR) computerised system, particularly in the context of limited formal guidelines and the known subjectivity of antepartum assessments. The findings are intended to support improvements in training and the standardization of CTG interpretation. MethodsFive senior clinicians with expertise in fetal monitoring independently annotated 105 15-minute fetal heart rate (FHR) traces using structured web-based software. For each trace, participants identified the baseline, accelerations and decelerations and categorised variability. Inter-observer agreement was assessed using intraclass correlation (ICC) for baselines and Fleiss{kappa} for variability. Sensitivity and positive predictive value (PPV) for acceleration and deceleration detection were calculated relative to majority-voted results. The DR algorithm was used to identify the same patterns and the output was compared against the clinical consensus annotations. ResultsBaseline agreement was excellent among participants (ICC = 1.0, 95% CI 1.00- 1.00). Variability classifications showed only moderate concordance (Fleiss{kappa} = 0.39, 95% CI -0.01-0.56). Detection of accelerations and decelerations varied across clinicians (sensitivity 39.2-97.2%, PPV 39.7-91.3%). The DR system showed good agreement for accelerations (sensitivity = 64.8%, 95% CI 57.1-72.1%; PPV = 85.0%, 95% CI 79.5-89.8%) but poor agreement for decelerations (sensitivity = 50.0%, 95% CI 14.3-75.0%; PPV = 20.0%, 95% CI 4.2-39.1%). DR-classified variability showed minimal agreement with clinical ratings (Fleiss{kappa} = 0.002, 95% CI -0.007-0.027). ConclusionsAntepartum CTG interpretation remains inconsistent for identification of decelerations and variability. While baseline assessment appears robust, current clinical and algorithmic approaches show limited agreement for more ambiguous patterns. These findings support the need for updated training and refined algorithms to improve reliability in antepartum fetal surveillance.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.