Back

Comorbidity structure as an inductive bias: Comparing output-head designs for multi-label prediction of diabetes and myocardial infarction complications

Asumboya, W. A.; Agbenorhevi, P. K.; Adams, C. F.; Ayariga, D. A.; Adjadeh, T.; Adams Ziblim, S.; Kwofie, S. K.

2026-06-23 bioinformatics
10.64898/2026.06.18.733068 bioRxiv
Show abstract

BackgroundClinical complications are often predicted with separate sigmoid outputs, even when the target labels arise from related pathophysiological processes. This paper asks whether output-layer choice should reflect both predictive convenience and the biological structure assumed among complications. The central premise is that label-dependence mechanisms are explicit hypotheses about comorbidity, not generic modelling additions. MethodsOutput-head assumptions were compared across two clinically distinct multi-label prediction tasks. In Type 2 diabetes (T2D), six heads were evaluated for nephropathy, neuropathy, and retinopathy: independent baseline, linear additive, multiplicative, symmetric conditional random field (CRF), residual multilayer perceptron (MLP), and combined additive-multiplicative. In myocardial infarction (MI), four heads were evaluated for ventricular tachycardia, ventricular fibrillation, and atrioventricular block: independent baseline, linear additive, multiplicative, and symmetric CRF. All experiments used five training data fractions and seven independent seeds, with the same shared-backbone protocol within each disease setting. ResultsIn T2D, the symmetric CRF gave the most consistent improvement pattern, ranking highest at full data and at the two lowest data fractions while adding only three interaction parameters. At 20% training data, it was the only interaction head whose aggregate mean exceeded the independent baseline. The residual MLP, despite 123 interaction parameters, remained below the baseline across all T2D fractions. In MI, rankings changed across fractions: the multiplicative head led at 80% and 60%, the CRF led at 100% and 20%, and the baseline led at 40%. The combined additive-multiplicative head did not improve robustness in T2D and showed the largest negative baseline-relative deviations at lower fractions. ConclusionThe findings support a biology-guided view of output-layer design. A small constrained mechanism was most useful when its symmetry matched the shared microvascular structure of T2D, whereas the heterogeneous electrophysiology of MI produced no stable winner. Output-layer choice should therefore be reported and defended as an assumption about disease structure instead of a routine hyperparameter decision. Author summaryMany clinical prediction models treat complications as separate outcomes, even when clinicians know they often arise together. We studied whether the last layer of a model should reflect that biological knowledge. We compared several output heads across two disease settings: Type 2 diabetes, where nephropathy, neuropathy, and retinopathy share a common microvascular origin, and myocardial infarction, where electrical complications arise from a mixture of shared and location-specific mechanisms. We found that a small symmetric CRF head was most useful in the diabetes task, especially when training data were limited, while no single interaction head dominated in myocardial infarction. This suggests that modelling comorbidity is not only a technical choice; it is a statement about how disease processes relate to one another. Our results encourage researchers to report and justify output-layer design as part of the clinical modelling argument, rather than treating it as a routine hyperparameter.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 0.5%
29.5%
2
Scientific Reports
3612 papers in training set
Top 9%
7.0%
3
BioData Mining
22 papers in training set
Top 0.1%
5.7%
4
PLOS ONE
5266 papers in training set
Top 27%
5.7%
5
Computers in Biology and Medicine
128 papers in training set
Top 0.7%
4.5%
50% of probability mass above
6
Journal of Translational Medicine
57 papers in training set
Top 0.3%
2.7%
7
eBioMedicine
183 papers in training set
Top 2%
2.2%
8
Diabetologia
44 papers in training set
Top 0.4%
2.0%
9
Journal of the American Medical Informatics Association
71 papers in training set
Top 1%
2.0%
10
Bioinformatics
1204 papers in training set
Top 7%
1.8%
11
npj Digital Medicine
118 papers in training set
Top 2%
1.5%
12
Brain Communications
166 papers in training set
Top 2%
1.5%
13
Nature Communications
5641 papers in training set
Top 48%
1.4%
14
eClinicalMedicine
77 papers in training set
Top 1%
1.4%
15
JAMIA Open
42 papers in training set
Top 1%
1.2%
16
Bioinformatics Advances
203 papers in training set
Top 4%
1.2%
17
Communications Medicine
113 papers in training set
Top 3%
1.2%
18
iScience
1154 papers in training set
Top 28%
1.1%
19
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
20
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
0.9%
21
Frontiers in Bioinformatics
49 papers in training set
Top 1%
0.9%
22
BMC Nephrology
18 papers in training set
Top 0.3%
0.9%
23
European Heart Journal
22 papers in training set
Top 1%
0.9%
24
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%
25
Expert Systems with Applications
11 papers in training set
Top 0.5%
0.6%
26
JMIR Public Health and Surveillance
45 papers in training set
Top 2%
0.6%
27
European Heart Journal - Digital Health
18 papers in training set
Top 1%
0.6%
28
PeerJ
308 papers in training set
Top 12%
0.6%
29
Circulation: Heart Failure
14 papers in training set
Top 0.6%
0.6%
30
GENETICS
483 papers in training set
Top 5%
0.6%