An End-to-End Two-Stream Network Based on RGB Flow and Representation Flow for Human Action Recognition

With the rapid development of deep learning, the performance of computer vision tasks has significantly improved, making the two-stream neural network model a popular research topic for video-based action recognition. Traditional models use RGB and optical-flow streams, which offer promising performance but at high computational cost and complexity. To address these issues, we introduce a representation flow algorithm and replace the optical-flow branch in the egocentric action recognition model with this new branch for end-to-end training. This can greatly reduce computational cost and prediction runtime. Our new model, applied to egocentric action recognition, incorporates class attention maps (CAMs) to enhance recognition accuracy and uses ConvLSTM for spatio-temporal encoding with spatial attention. Evaluated on the GTEA61, EGTEA GAZE+, and HMDB datasets, our model matches the original model’s accuracy on GTEA61 and exceeds it by 0.65% and 0.84% on EGTEA GAZE+ and HMDB, respectively. The proposed model’s runtimes for predicting are significantly faster: 0.1881s, 0.1503s, and 0.1459s compared to the original’s 101.6795s, 25.3799s, and 203.9958s. We also conducted ablation studies to analyze the impact of different parameters on model performance.

Paper

References (16)

Scroll for more · 4 remaining

Similar papers

© 2026 NYSGPT2525 LLC