Back

Limitations of Public Biomechanical Movement Datasets for Deep Learning: Issues of Metadata, Standardization, and Variety in Motion Types

Friemert, D.; Schnur, D.; Runkel, S.; Borsch, J.; Karamanidis, K.; Dellen, B.; Thieme, L.; Fiedler, A.; Jaeckel, U.; Hartmann, U.

2025-05-29 sports medicine
10.1101/2025.05.29.25328474 medRxiv
Show abstract

Biomechanical data play a crucial role in health research by providing relevant information about musculoskeletal mechanical loading and neuromuscular dysfunction during movement, which are essential for optimizing treatment strategies and improving patient outcomes. The relevance of such data for decision making processes in clinical settings depends on its quality and volume, as larger datasets enable more robust conclusions across diverse populations and conditions. In machine learning, extensive and representative datasets are vital for training algorithms to develop accurate predictive models. This systematic literature review explores the current landscape of biomechanical datasets and databases, examining existing resources, identifying gaps in data availability and quality, and discussing the potential for a new, structured database to address these shortcomings. We conducted a comprehensive search across PubMed, IEEE Xplore, Scopus, ScienceDirect, Springer, and arXiv using keywords related to biomechanics, gait, and human activity recognition (HAR). The review identified studies covering various motion types, data formats, and metadata availability. Our findings highlight significant challenges, including inconsistent metadata, lack of standardization during data acquisition and data processing, limited motion variety, and accessibility issues. Addressing these challenges through standardized data formats, comprehensive metadata, and enhanced accessibility could greatly advance decision making processes via machine learning approaches for clinical and therapeutic settings. This review emphasizes the need for improved biomechanical data resources to support the development of accurate machine learning models and innovative clinical interventions.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.