MotionScope AI: Comprehensive Human Activity Recognition Through Integrated Pose Analysis and Temporal Modeling
This paper introduces a practical deep learning framework for recognizing human activities in indoor environments using MediaPipe pose estimation with multi-branch bidirectional LSTM architecture. We extract comprehensive features—including 3D pose landmarks, hand gestures, velocity, acceleration, and joint angles—resulting in 685-dimensional vectors per frame. Our multi-branch design processes each feature type through specialized BiLSTM pathways with attention mechanisms, enabling the model to learn distinct spatial, temporal, and structural patterns. To ensure robust performance in real-world scenarios, we incorporate label smoothing, gradient clipping, and adaptive learning strategies. Evaluated on the IndoorActionDataset with 8 activity classes, our approach achieves 95.6% test accuracy, significantly outperforming OpenPose-based methods (87.2%) and 3D-CNN with Transformer architectures (75.1%). With only 23ms inference latency, the system demonstrates practical viability for real-time deployment on resource-constrained devices. The results confirm that thoughtful feature engineering combined with attention-driven temporal modeling can deliver both accuracy and efficiency for activity recognition tasks.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex