Back

A prevalence-incidence-clearance model for interval-censored screening and surveillance data in a population with an elevated disease risk at baseline

Kroon, K. R.; Bogaards, J. A.; Meijer, C. J.; Berkhof, J.

2025-04-20 infectious diseases
10.1101/2025.04.19.25325991 medRxiv
Show abstract

Accurate risk assessment is essential for screening and surveillance programs, but this is complicated when considering baseline conditions linked to an increased risk of disease that may decline over time (e.g., certain infections and viral-induced disease). In longitudinal screening and surveillance studies, individuals may have prevalent disease at baseline or develop it during follow-up, either from their baseline condition ("early" event) or from a new condition ("late" event). Additionally, data are interval-censored between visits, making the exact time of disease onset unknown. We propose a prevalence-incidence-clearance model for interval-censored data to estimate cumulative disease risk based on individual risk factors, with the motivating example of human papillomavirus (HPV) infections, which may progress to high-grade cervical lesions and cancer (CIN2+). Early events are modelled with an exponential competing risks framework, where HPV infections either progress to CIN2+ or to a (latent) "clearance" state. Late events are modelled by adding a background risk. Parameters are estimated with an expectation-maximisation algorithm with weakly informative Cauchy priors. The algorithm was validated through simulation studies and applied to screening and post-treatment surveillance data from the Netherlands. Our model accurately predicts cumulative CIN2+ risk in HPV-positive women and fits the observed cumulative incidence curve better than existing methods. Furthermore, it provides easily interpretable parameters and its baseline hazard can be checked for lack of fit. This is especially important when applying the model to facilitate decision-making for national programs.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.