Data Diversity vs. Model Complexity in the Prediction of Pediatric Bipolar Disorder: Evidence from Academic and Community Clinical Samples
Shi, Z.; Youngstrom, E. A.; Liu, Y.; Youngstrom, J. K.; Findling, R. L.
Show abstract
Pediatric bipolar disorder is challenging to diagnose accurately due to symptom heterogeneity. More standardized and data-driven approaches are needed to enhance diagnostic reliability. We evaluated a clinical decision tool (nomogram), statistical methods (logistic regression, LASSO), machine learning (support vector machine, random forest, k-nearest neighbors, extreme gradient boosting), and deep learning model (multilayer perceptron) for pediatric bipolar disorder prediction across two datasets collected in academic (N=550) and community (N=511) clinical settings. We compared three modeling strategies: cross-dataset validation, cross-dataset with interaction terms, and mixed-dataset. We assessed model performance using discrimination ability, calibration, and predictor importance ranking. In the baseline cross-dataset approach, all models showed good internal discrimination in the academic dataset; but external discrimination in the community dataset substantially declined. Interaction-enhanced models slightly improved internal discrimination but not external performance or calibration. Recalibration prominently improved cross-dataset calibration without compromising discrimination, indicating that transportability problems were largely driven by probability scaling. Models trained on mixed datasets exhibited much stronger external discrimination and calibration. Across models and training strategies, family risk and PGBI-10M were consistently ranked as the most important predictors. Predictive models for pediatric bipolar disorder showed strong internal performance but limited cross-setting generalizability due to dataset shift and miscalibration. Increasing model complexity did not improve external performance, whereas training on pooled data substantially improved both discrimination and calibration. Findings suggest that sampling diversity, rather than model complexity, is more valuable for developing clinically useful and generalizable psychiatric prediction models, underscoring the importance of open and collaborative datasets.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Prediction of Adolescent Internalizing Disorder Risk: Evidence from the Norwegian Mother, Father, and Child Cohort Study 93%
- Depression is Associated with Treatment Response Trajectories in Adults with Prolonged Grief Disorder: A Machine Learning Analysis 93%
- Exploring the Efficacy and Potential of Large Language Models for Depression: A Systematic Review 92%
Similar papers in this journal
- Differential Treatment Benefit Prediction For Treatment Selection in Depression: A Deep Learning Analysis of STAR*D and CO-MED Data 94%
- Computational Mechanisms of Approach-Avoidance Conflict Predictively Differentiate Between Affective and Substance Use Disorders 94%
- Classifying Obsessive-Compulsive Disorder from Resting-State EEG using Convolutional Neural Networks: A Pilot Study 91%
Similar papers in this journal
- Machine Learning Models Predict the Emergence of Depression in Argentinean College Students during Periods of COVID-19 Quarantine 95%
- Applications of Large Language Models in Psychiatry: A Systematic Review 92%
- Understanding Psychiatric Illness Through Natural Language Processing (UNDERPIN): Rationale, Design, and Methodology 90%
Similar papers in this journal
- Predicting involuntary admission following inpatient psychiatric treatment using machine learning trained on electronic health record data 93%
- The Broad Structure of Psychopathology in the All of Us Research Program 91%
- A Comparison of Pruning During Multi-Step Planning in Depressed and Healthy Individuals 91%
Similar papers in this journal
- The path toward generalizable clinical prediction models 93%
- Psychosis Prognosis Predictor: A Continuous and Uncertainty-Aware Prediction of Treatment Outcome in First-Episode Psychosis 92%
- Receiving information on machine learning-based clinical decision support systems in psychiatric services increases staff trust in these systems: A randomized survey experiment 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.