Back

Handling Missing Data in Participants with a Baseline but No Post-baseline Data

Mallinckrodt, C. H.; Lipkovich, I.; Dickson, S.; Hendrix, S. B.; Molenberghs, G.

2025-03-26 pharmacology and therapeutics
10.1101/2025.03.26.25324679 medRxiv
Show abstract

Clinical trial participants who are randomized to treatment and have a baseline value but no post-baseline data pose a unique challenge. These patients need to be included to preserve randomization. However, there is no information about the outcome or the intercurrent events that led to missing data. Composite and hypothetical strategies can accommodate participants with no post-baseline data. The present study investigated common methods based on hypothetical strategies in example data sets and in a simulation study. Because there is no information about the treatment effect when there is no post-baseline data, it is not surprising that various models yielded similar results. In simulated data, treatment contrasts were not biased when the reason for missing all post-baseline data was random, treatment related, or outcome related. Bias occurred only when missingness was treatment and outcome related, and arose from a missing not at random mechanism. The bias was not large. The maximum Type I error rate across all scenarios and methods was 6.4%. Imputing no change for missing data at the first postbaseline visit and assuming missing at random for all other missing values yielded a maximum Type I error rate of 5.1%. Given the idiosyncratic nature of clinical trials, no universally best analytic approach exists for dealing with participants that have a baseline but no post-baseline data. Analysts can choose among these methods to provide an approach tailored to the situation at hand.

Published in Pharmaceutical Statistics · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.