Back

Predicting long-term adverse outcomes after neonatal intensive care

Ogretir, M.; Kaipainen, V.; Leskinen, M.; Lahdesmaki, H.; Koskinen, M.

2026-03-31 pediatrics
10.64898/2026.03.26.26348580 medRxiv
Show abstract

Neonates requiring intensive care are at increased risk for long-term neuropsychiatric disorders. However, clinical adoption of risk prediction models remains limited when their performance lacks adequate interpretability for informed clinical decision-making. Here, we investigated whether longitudinal neonatal electronic health record (EHR) data from the first 90 days of life can support clinically meaningful interpretation of long-term risk signals for major neuropsychiatric diagnoses by age seven. In a retrospective register-based cohort of 17,655 at-risk children from an academic medical center, of whom 8.0\% (1,420) received a major neuropsychiatric diagnosis during follow-up, we applied a time-aware transformer model (Self-supervised Transformer for Time-Series; STraTS) and thoroughly evaluated its predictions using three complementary interpretability approaches: perturbation-based variable importance, value-dependent effect analysis, and leave-one-out (LOO) feature attribution. STraTS achieved the highest area under the precision--recall curve (AUPRC 0.171 {+/-} 0.022), compared with Random Forest (0.166 {+/-} 0.008), logistic regression (0.151 {+/-} 0.007), and XGBoost (0.128 {+/-} 0.010). Across interpretability methods, five predictors were consistently identified: birth weight, gender, Apgar score at 1 minute, umbilical serum thyroid stimulating hormone (uS-TSH), and treatment time in hospital. Indicators of early clinical severity, including chromosomal abnormalities and neonatal cerebral-status disturbances, showed the largest risk-increasing effects. Furthermore, the model's learned vector representations of subject-specific EHR sequences formed clinically coherent latent embeddings that reflect population heterogeneity along established perinatal risk dimensions. These findings demonstrate that combining multiple complementary interpretability methods yields stable, clinically plausible risk signals while revealing limitations that would remain undetected by any single approach, highlighting the importance of careful interpretability analysis of deep learning-based risk predictions.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.3%
18.3%
2
BioData Mining
22 papers in training set
Top 0.1%
9.6%
3
Communications Medicine
113 papers in training set
Top 0.2%
7.8%
4
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 0.1%
6.7%
5
Scientific Reports
3612 papers in training set
Top 13%
6.2%
6
Nature Medicine
125 papers in training set
Top 0.5%
4.0%
50% of probability mass above
7
Archives of Disease in Childhood
16 papers in training set
Top 0.1%
3.4%
8
The Journal of Pediatrics
16 papers in training set
Top 0.1%
3.4%
9
Pediatric Research
21 papers in training set
Top 0.1%
3.2%
10
JAMA Network Open
130 papers in training set
Top 1%
2.7%
11
BMC Medicine
176 papers in training set
Top 2%
2.4%
12
Translational Psychiatry
260 papers in training set
Top 2%
2.4%
13
Cell Reports Medicine
153 papers in training set
Top 1%
2.4%
14
PLOS ONE
5266 papers in training set
Top 45%
2.1%
15
eBioMedicine
183 papers in training set
Top 3%
1.7%
16
Human Genetics and Genomics Advances
84 papers in training set
Top 1%
1.7%
17
Developmental Cognitive Neuroscience
96 papers in training set
Top 0.7%
1.4%
18
Science Translational Medicine
127 papers in training set
Top 2%
1.1%
19
Genome Medicine
183 papers in training set
Top 4%
1.1%
20
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 35%
1.1%
21
Frontiers in Pediatrics
32 papers in training set
Top 0.8%
1.1%
22
PLOS Medicine
110 papers in training set
Top 3%
1.0%
23
Psychological Medicine
88 papers in training set
Top 2%
1.0%
24
JAMA Pediatrics
10 papers in training set
Top 0.2%
0.8%
25
Nature Communications
5641 papers in training set
Top 57%
0.8%
26
GigaScience
212 papers in training set
Top 5%
0.6%
27
Journal of Clinical Epidemiology
31 papers in training set
Top 1.0%
0.6%
28
Journal of Allergy and Clinical Immunology
27 papers in training set
Top 0.6%
0.6%