Automatic sleep staging in patients with suspected sleep disorders: a comparison of existing methods on portable setups
Gunter, K. M.; Dorier, A.; Bowring, F.; Dennis, G.; Lo, C.; Quinnell, T.; Symmonds, M.; Ratti, P.-L.; Hu, M. T.; Villarroel, M.
Show abstract
Background: Automatic sleep staging algorithms are increasingly applied in clinical and home-based recordings. However, their performance may degrade when transferred to new montages and clinical populations. This is particularly relevant in reduced-channel portable PSG and in disorders such as REM sleep behaviour disorder (RBD), where altered sleep architecture may challenge pretrained models. Objective: To evaluate and compare multiple open-source sleep staging algorithms on a minimal portable PSG setup in controls and patients with and without RBD, and to assess the impact of fine-tuning on clinic-ascertained data. Methods: Six open-source models were applied to 76 subjects recruited from three clinical sleep medicine sites. Performance was assessed using accuracy, F1 scores, and Cohen's kappa, both overall and per sleep stage. Each model was evaluated out-of-the-box and after fine-tuning on clinical data. Results: Out-of-the-box performance varied substantially across models (Cohen's kappa 0.21-0.54). Fine-tuning consistently improved agreement, with the best-performing model (GSSC) reaching Cohen's kappa = 0.58 indicating moderate to good agreement. Performance was highest in controls and lower in patient groups. N3 was the most reliably classified stage across models, whereas N1 remained consistently challenging. REM classification improved after fine-tuning in several architectures but remained model, and subgroup-dependent, particularly in RBD subjects. Conclusion: Fine-tuning substantially mitigates domain shift, updating model parameters to align with new data distributions, when applying automatic sleep staging algorithms to portable clinical recordings. Model architecture influences robustness, with feature-learning approaches demonstrating greater adaptability than fixed-feature models. Despite moderate agreement after adaptation, performance, especially for REM and N1 remains insufficient for fully automated diagnostic use in clinical populations.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Novel Digital Markers of Sleep Dynamics: A Causal Inference Approach Revealing Age and Gender Phenotypes in Obstructive Sleep Apnea 95%
- Topographical relocation of adolescent sleep spindles reveals a new maturational pattern of the human brain 95%
- Is sleep apnea-hypopnea index relevant for impaired brain perfusion and desaturation in patients with severe obstructive sleep apnea syndromes? 94%
Similar papers in this journal
- A foundational transformer leveraging full night, multichannel sleep study data accurately classifies sleep stages 98%
- Evaluation of Dreem headband for sleep staging and EEG spectral analysis in people living with Alzheimer’s and older adults 97%
- Corticothalamic modelling of sleep neurophysiology with applications to mobile EEG 95%
Similar papers in this journal
- On the development of sleep states in the first weeks of life 97%
- Discrimination of sleep and wake periods from a hip-worn raw acceleration sensor using recurrent neural networks 95%
- Home-EEG assessment of possible compensatory mechanisms for sleep disruption in highly irregular shift workers - The ANCHOR study 94%
Similar papers in this journal
- The potential of ensemble-based automated sleep staging on single-channel EEG signal from a wearable device 97%
- Looking for a reference for large datasets: relative reliability of visual and automatic sleep scoring 96%
- Automated real-time EEG sleep spindle detection for brain state-dependent brain stimulation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.