Generalizable CNN-based detection of form and speed errors in physiotherapy using multimodal activity images
Accurate evaluation of physiotherapy exercises is critical to effective rehabilitation and injury prevention. We present a multimodal deep learning framework for detecting both form and speed errors in shoulder rehabilitation exercises, using data from a standard RGB camera and a wrist-worn smart band. The exteroceptive vision data are first transformed into time-series signals and fused with proprioceptive accelerometer data using a signal-to-activity imaging algorithm that captures both spatial and temporal dynamics. To enhance model performance and interpretability, the framework combines convolutional neural networks (CNNs) with computed statistical features (CF), forming a hybrid CNN+CF architecture. This integration significantly improves accuracy—particularly under leave-one-out cross-validation (LOO-CV)—ensuring robust generalization to unseen users, a critical requirement for real-world deployment. The system achieves 99% accuracy in half-half validation and 75% in LOO-CV across 12 exercise classes. A hierarchical classification strategy further improves performance, achieving 95% accuracy for top-level categories and averaging 85% for detailed detection of form and speed errors. Additionally, Layer-wise Relevance Propagation (LRP) is employed to interpret model decisions and highlight contributing features. The proposed CNN+CF model, underpinned by signal imaging and trained with limited data, serves as a proof of concept and demonstrates strong robustness to sensor noise and inter-subject variability, offering a scalable solution for unsupervised home-based rehabilitation monitoring.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex