Development and evaluation of wrist- and thigh-worn accelerometer algorithms using self-training machine learning models for classification of activity type and posture: towards device placement-agnostic methods in the ProPASS consortium
Ahmadi, M.; Koemel, N. A.; Biswas, R.; Holtermann, A.; Koster, A.; Atkin, A.; Pulsford, R.; del Pozo Cruz, B.; Rangul, V.; Mitchell, J. J.; Blodgett, J. M.; Hettiarachchi, P.; Svartengren, M.; Granat, M.; Clark, B.; Hamer, M.; Stamatakis, E.
Show abstract
BackgroundWearable accelerometers are widely used in health research, but differing placements (e.g. wrist vs. thigh) hinder harmonising activity classification across studies. Prior studies report 1.5- to 2.0-fold differences in physical activity level by wear location, hampering data comparability, and compromising the potential for pooling data to develop consortia and carrying out Meta- and Individual Participant Data analysis. Although supervised machine learning is increasingly used in wearables research, its reliance on extensive labelled data limits its use in free-living datasets. Semi-supervised learning offers an efficient alternative by using laboratory collected labelled data to iteratively self-train models on unlabelled free-living data. Using a self-training approach, the aim of this study was to train and evaluate algorithms for wrist- and thigh-worn devices to facilitate harmonisation of posture and activity type classification between placements. MethodsA total of 146 participants aged 30-75 years completed either structured laboratory-based activity trials or one of two independent free-living assessments while wearing Axivity AX3 accelerometers on the wrist and thigh. For each placement, a supervised Random Forest classifier was initially trained using a labelled laboratory dataset (n=40) to classify sitting, standing, walking, running, stair climbing, and cycling - and then re-trained using self-training on free-living data (n=53, independent to the laboratory study sample). The final models were validated using another hold-out free-living independent dataset (n=53) with ground-truth activity labels obtained via direct video observation. Overall model comparison and performance was assessed using accuracy, kappa statistic, and F1 scores. Individual activity class comparison and performance was evaluated using equivalence testing, confusion matrices, and coefficient of variation between the wrist and thigh estimates. ResultsDuring a total of 43,800 minutes, of which 19,080 minutes were in the hold-out dataset, both self-trained models achieved high overall classification accuracy: 91.8% (SD = 6.8%) for the wrist and 95.1% (SD = 5.4%) for the thigh. The overall F1 score was 88.2 (SD = 9.6%) for the wrist classifier and 90.1 (SD = 9.3%) for the thigh classifier. Equivalence testing demonstrated that both classifiers produced activity duration estimates statistically equivalent to ground-truth for all activity types except stair climbing. Confusion matrices for the wrist demonstrated very good to excellent (88% - 97%) classification accuracy for sitting, walking, running, and cycling, and good accuracy for standing and stair climbing (71%-78%). For the thigh, classification performance was very good to excellent (83% - 98%) across sitting, standing, walking, running, and cycling, with good accuracy for stair climbing (75%). The coefficient of variation values ranged from 0.022 for running to 0.140 for standing. ConclusionThese findings highlight the potential of self-training models to support harmonisation of wearable accelerometer data collected using different wear placements in ProPASS and other consortia. Self-training models reduce reliance on extensive labelled data and demonstrated high activity type classification accuracy for both wrist- and thigh-worn accelerometers, with a high degree of agreement and equivalence with ground-truth data across almost all activity types.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Behaviour-based movement cut-off points in 3-year old children comparing wrist- with hip-worn actigraphs MW8 and GT3X 95%
- Comparison of raw accelerometry data from ActiGraph, Apple Watch, Garmin, and Fitbit using a mechanical shaker table 95%
- From Movement to METs: A Validation of ActTrust(R) for Energy Expenditure Estimation and Physical Activity Classification in Young Adults 95%
Similar papers in this journal
- Vigorous intermittent lifestyle physical activity (VILPA) and mortality risk among US adults: a wearables-based national cohort study 93%
- Descriptive epidemiology of physical activity energy expenditure in UK adults. The Fenland Study. 92%
- Development and Usability of a Mobile Ecological Momentary Assessment Platform for Dietary Surveillance in the U.S. 91%
Similar papers in this journal
- Passive Detection of COVID-19 with Wearable Sensors and Explainable Machine Learning Algorithms 94%
- An aging focused unobtrusive and Privacy-Preserving Digital Behaviorome 93%
- Digital health technologies and machine learning augment patient reported outcomes to remotely characterise rheumatoid arthritis 92%
Similar papers in this journal
- Ambulatory physiological measures obtained under naturalistic urban mobility conditions have acceptable reliability 94%
- An infection prediction model developed from inpatient data can predict out-of-hospital COVID-19 infections from wearable data when controlled for dataset shift 93%
- Sociodemographic Characteristics of Missing Data in Digital Phenotyping 93%
Similar papers in this journal
- Cohort Profile: Baseline characteristics and design of the McMaster Monitoring My Mobility (MacM3) Study, a prospective digital mobility cohort of community-dwelling older Canadians from Southern Ontario 94%
- Wearable Sensing in Eating Episode Monitoring: An Updated Systematic Review Protocol 92%
- Correlates of and changes in aerobic physical activity and strength training before and after the onset of COVID-19 pandemic in the UK – findings from the HEBECO study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.