Back

DACr - An Algorithm for Treating Diverse Missing Values in Large Data, with Application to Heart Transplantation

Dai, W.; Sun, J.; Xu, J.

2026-01-23 health informatics
10.64898/2026.01.20.26343799 medRxiv
Show abstract

The increasing use of Real-World Data (RWD) in clinical research is critical for evidence-based decision making but presents challenges to data analytics. Unlike in Randomized Controlled Trials (RCTs), missingness in RWD occurs often and can include complex patterns which may or may not be Missing at Random (MAR). Informative absence (Missing Not at Random, or MNAR) occurs when the absence of data itself is a clinical signal. Applications of ad-hoc methods, or even popular "universal" methods, can lead to biased inferences when these varied data problems are not appropriately addressed. This paper introduces a multi-step algorithm known as Divide And Conquer, rejoin (DACr, pronounced as "DACK-er") for handling missing data in RWD. We illustrate the algorithms application using a large-scale, high-dimensional heart transplant dataset (from p = 1107 to 527 after initial screening) from the Scientific Registry of Transplant Recipients (SRTR). This "divide and conquer" idea handles variables based on data type, the proportion, and type of missingness. It also allows for the targeted application of justified data handling techniques: appropriate imputation method is used for variables assumed to be MAR, while variables with plausible MNAR mechanisms are recoded as categorical with an extra level of "missing" to preserve their informative signal. The proposed algorithm provides a systematic treatment for practical researchers.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.