Back

Development and Internal Validation of Models Predicting the Health Insurance Status of Participants in the German National Cohort

Hrudey, I.; Swart, E.; Baurecht, H.; Becher, H.; Damms-Machado, A.; Hoffmann, W.; Jöckel, K.-H.; Kartschmit, N.; Katzke, V.; Keil, T.; Kollhorst, B.; Leitzmann, M.; Meinke-Franze, C.; Michels, K. B.; Mikolajczyk, R.; Niedermaier, T.; Pigeot, I.; Schipf, S.; Schmidt, B.; Walter, B.; Willich, S.; Wolff, R.; Stallmann, C.

2024-04-12 epidemiology
10.1101/2024.04.09.24305544 medRxiv
Show abstract

BackgroundIn Germany, all citizens must purchase health insurance, in either statutory (SHI) or private health insurance (PHI). Because of the division into SHI and PHI, person insurances status is an important variable for studies in the context of public health research. In the German National Cohort (NAKO), the variable on self-reported health insurance status of the participants has a high proportion of missing values (55.4%). The aim of our study was to develop and internally validate models to predict the health insurance status of NAKO baseline survey participants in order to replace missing values. In this respect, our research interest was focused on the question to which extent socio-demographic characteristics are suitable for predicting health insurance status. MethodsWe developed two prediction models including 53,796 participants to estimate the probability that a participant is either member of a SHI (model 1) or PHI (model 2). We identified eight predictors by literature research: occupation, income, education, sex, age, employment status, residential area, and marital status. The predictive performance was determined in the internal validation considering discrimination and calibration. Discrimination was assessed based on the Area Under the Curve (AUC) and the Receiver Operating Characteristic (ROC) curve and calibration was assessed based on the calibration slope and calibration plot. ResultsIn model 1, the AUC was 0.91 (95% CI: 0.91-0.92) and the calibration slope was 0.97 (95% CI: 0.97-0.97). Model 2 had an AUC of 0.91 (95% CI: 0.90-0.91) and a calibration slope of 0.97 (95% CI: 0.97-0.97). Based on the calculated performance parameters both models turned out to show an almost ideal discrimination and calibration. Employment status and household income and to a lesser extent educational level, age, sex, marital status, and residential area are suitable for predicting health insurance status. ConclusionsSocio-demographic characteristics especially employment status and household income assessed at NAKOs baseline were suitable for predicting the statutory and private health insurance status. However, before applying the prediction models in other studies, an external validation in population-based studies is recommended.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.