Back

Choice of estimands and estimators affected the interpretation of results for some outcomes in a cluster-randomised trial (RESTORE) due to informative cluster size

Bi, D.; Copas, A.; Li, F.; Harhay, M. O.; Kahan, B. C.

2026-05-06 intensive care and critical care medicine
10.64898/2026.05.05.26352371 medRxiv
Show abstract

Background and objectiveIn cluster-randomised trials (CRTs), different estimands can be targeted, such as the individual-or cluster-average effect. These two estimands can differ in magnitude when outcomes or treatment effects vary with cluster size (termed informative cluster size). When informative cluster size is present, commonly used estimators for CRTs, such as mixed-effects model and generalised estimating equations with an exchangeable correlation structure (termed GEEs(exch)), can be biased for both these estimands. With little documented evaluation of When informative cluster size, it is currently unknown how commonly it occurs in practice. The aim of this work was to explore whether informative cluster size is present in a published CRT and to investigate its impact on trial results. MethodsWe re-analysed the RESTORE CRT, which compared protocolised sedation with usual care for critically ill children. For each outcome, we first modelled the association between cluster size and outcome/treatment effect; next, we assessed the impact of informative cluster size by comparing differences between (i) individual-vs. cluster-average estimates and (ii) estimates from mixed-effects models and GEEs(exch) (which can be affected by informative cluster size) to those from IEEs (which are robust to informative cluster size). ResultsWe found evidence of an association between cluster size and either outcomes or treatment effects for 16/33 outcomes (48%). This led to statistically significant differences between the individual- and cluster-average treatment effects for 5 of 33 outcomes (15%). There were >10% differences between (i) individual- and cluster-average treatment effect estimates for 17 outcomes (52%) and (ii) estimates from mixed-effects models/GEEs(exch) and estimates from unweighted IEEs for 13 outcomes (39%). For some outcomes, differences in the choice of estimator or estimand led to differences in the interpretation of results. For example, for the outcome postextubation stridor, the individual-average estimate showed a significant harmful effect (OR=1.65, 95% CI 1.02 to 2.67), unlike the cluster-average (OR=1.38, 95% CI 0.87 to 2.19) or GEEs(exch) estimate (OR=1.57, 95% CI 0.98, 2.50). Discussioninformative cluster size can occur in CRTs, and the use of estimators that are not clearly aligned to the target estimand can affect the interpretation of some results. What is new?O_ST_ABSKey findingsC_ST_ABSO_LIThis re-analysis of the RESTORE cluster randomised trial found that choice of estimand and estimator could affect the interpretation of results for some outcomes C_LI What this adds to what is knownO_LIThis work provides empirical evidence that informative cluster size can occur in cluster randomised trials, and can affect results based on the choice of estimand or estimator C_LI What is the implication and what should we change nowO_LITrialists should clearly define their target estimand and choose an estimator that is aligned to that estimand C_LIO_LICareful consideration of the plausibility of assumptions underpinning each estimator, including the likelihood of informative cluster size, can help ensure appropriate analysis methods are used C_LIO_LIWhen mixed-effects models or GEEs with an exchangeable correlation structure are used, sensitivity analyses using independence estimating equations or other appropriate methods should be used to evaluate the robustness of results to informative cluster size C_LI

Published in Journal of Clinical Epidemiology (predicted rank #11) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.