The Sleep-Wake Classification Performance of Pediatric-Trained Machine Learning Algorithms for Raw Accelerometer Data
Chen, P.-W.; Cielo, C.; Walsh, O.; Mcdonald, M.; Song, P. X.; Goldstein, C.; Moreno, J. P.; Jansen, E.; Mitchell, J. A.
Show abstract
Introduction: Actigraphy sleep-wake classification methods increasingly seek to leverage raw acceleration data and machine-learning-based classification, but performance evaluation in pediatrics is limited. We trained machine-learning models using pediatric data and compared their sleep-wake classification performance with existing algorithms for children. Methods: Sixty-five children (46% female, ages 5.3 to 17.7 years) completed in-lab overnight polysomnography and wore a GENEActiv device on their non-dominant wrist. The acceleration data were converted into 30-second epochs and aligned with physician-scored sleep-wake data from electroencephalography. Seven machine-learning models were trained using leave-one-subject-out cross-validation. Epoch-by-epoch analyses generated performance metrics (e.g., balanced accuracy [BA]) and discrepancy analyses provided overall sleep duration bias estimates. The combination of highest performance and least bias was used to rank using Euclidean distance scores - where a lower score represents closer to perfect performance and zero bias. For benchmarking, we included GGIR sleep scoring algorithms and an adult trained random forest classifier. Results: Overall, 560.1 hours of polysomnography and actigraphy data were collected (74.4% of epochs were scored as sleep). The pediatric-trained local-global long-short term memory (LSTM) classifier had the most optimal epoch-by-epoch performance (e.g., BA=0.85, sensitivity=0.88, specificity=0.83, ROC-AUC=0.95, and Cohen kappa=0.67). These metrics exceeded that of an adult-trained random forest classifier and GGIR-based algorithms. Discrepancy analyses revealed that overall sleep duration was underestimated by an average of 25 minutes using the LSTM classifier with no proportional bias. Conclusion: We trained seven pediatric sleep-wake classifiers that had strong ability to detect sleep and wake, with the LSTM classifier being most optimal.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Methodological approach to sleep state misperception in insomnia disorder: comparison between multiple nights of actigraphy recordings and a single night of polysomnography recording 96%
- Application of Down-Phase Targeted Auditory Stimulation During Sleep in a Home Setting: A Feasibility Study Across Seven Consecutive Nights 95%
- Effects of one-night partial sleep deprivation on perivascular space volume fraction: Findings from the Stockholm Sleepy Brain Study 94%
Similar papers in this journal
- The potential of ensemble-based automated sleep staging on single-channel EEG signal from a wearable device 98%
- Respiration-Triggered Olfactory Stimulation ReducesObstructive Sleep Apnea Symptoms Severity: A Prospective Pilot Study 96%
- Targeted memory reactivation during post-learning sleep does not enhance motor memory consolidation in older adults 94%
Similar papers in this journal
- Topographical relocation of adolescent sleep spindles reveals a new maturational pattern of the human brain 95%
- Novel Digital Markers of Sleep Dynamics: A Causal Inference Approach Revealing Age and Gender Phenotypes in Obstructive Sleep Apnea 95%
- Sleep Regularity Index as a Novel Indicator of Sleep Disturbance in Stroke Survivors: A Secondary Data Analysis 94%
Similar papers in this journal
- A foundational transformer leveraging full night, multichannel sleep study data accurately classifies sleep stages 96%
- Evaluation of Dreem headband for sleep staging and EEG spectral analysis in people living with Alzheimer’s and older adults 96%
- The Aging Slow Wave: A Shifting Amalgam of Distinct Slow Wave and Spindle Coupling Subtypes Define Slow Wave Sleep Across the Human Lifespan 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.