Back

Single-label and Multi-label Classification for Disease Recognition with Special Consideration of Comorbidities

Schmiegel, S.; Marchi, H.; Roechter, M.-H.; Rudwaleit, M.; Fuchs, C.

2025-12-31 health informatics
10.64898/2025.12.23.25342901 medRxiv
Show abstract

Certain diseases require rapid treatment to avoid long-term consequences for patients. However, they may be difficult to recognize, especially if the symptoms are ambiguous and compatible with multiple possible diagnoses. Completing all necessary examinations often takes time, thereby prolonging patient suffering. Data-driven approaches, such as single-label classification (SLC) and multi-label classification (MLC), can help accelerate the diagnostic process and improve accuracy. These two approaches differ primarily in the number of classes they allow a sample to belong to: SLC assumes the classes being mutually exclusive so that each sample belongs to exactly one class whereas MLC supposes the classes being mutually inclusive, i.e. a sample can belong to several classes or none, acknowledging the possibility of comorbidities. Comparing SLC and MLC allows us to investigate whether disease recognition benefits from considering comorbidities. In this context, we aim to provide a conceptual framing of (differences between) the two approaches in model formulation, decision spaces and handling of class imbalance. To empirically assess their performance, we conduct a case study applying SLC and MLC to data from chronic pain patients. Our analysis yields an ambiguous picture of whether incorporating comorbidities improve disease recognition. The suitability of SLC and MLC is determined by multiple factors, notably the dependency structure among diseases and between diseases and covariates, as well as by data characteristics such as class imbalance. This highlights the importance of considering the specific characteristics of the data when selecting an appropriate classification approach for disease recognition and beyond.

Published in BMC Medical Research Methodology (predicted rank #6) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.