Back

Imputing partial birth dates using day of the week

Johnson, C. Y.

2025-10-13 epidemiology
10.1101/2025.10.08.25337453 medRxiv
Show abstract

BackgroundIn deidentified data, exact dates are suppressed to maintain confidentiality of research participants. When the partial date includes only month and year, researchers who need exact dates must impute a day of the month. In some deidentified datasets, day of the week is also provided, but this variable is uncommonly incorporated into the imputation of partial dates. Our objective was to examine the extent to which misclassification is reduced by incorporating day of the week into partial date imputation. MethodsWe simulated a population of 594,677 people using the distribution of birthdays in England and Wales in 2024. We imputed birth dates using four methods: (1) first day of the month, (2) 15th of the month, (3) randomly selecting a day of the month, and (4) randomly selecting a day of the month conditional on day of the week. We quantified misclassification as the median number of days between the imputed and true birth date and as the cumulative percentage of the population whose imputed birth date fell within a given number of weeks of their true birth date. ResultsIncorporating day of the week reduced misclassification, with a median of 7 days between imputed and exact birth date compared to 8-15 for the other methods. For nearly a quarter of the population, their imputed birth date was their true birth date, compared to 3% in other methods. However, using the 15th day of the month was the best method to ensure that no misclassification was greater than 3 weeks. ConclusionIncorporating day of the week into random day selection reduced misclassification. This method is easily accomplished in standard statistical software.

Published in Paediatric and Perinatal Epidemiology · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.