VPN++: Rethinking Video-Pose embeddings for understanding Activities of Daily Living

Many attempts have been made towards combining RGB and 3D poses for the\nrecognition of Activities of Daily Living (ADL). ADL may look very similar and\noften necessitate to model fine-grained details to distinguish them. Because\nthe recent 3D ConvNets are too rigid to capture the subtle visual patterns\nacross an action, this research direction is dominated by methods combining RGB\nand 3D Poses. But the cost of computing 3D poses from RGB stream is high in the\nabsence of appropriate sensors. This limits the usage of aforementioned\napproaches in real-world applications requiring low latency. Then, how to best\ntake advantage of 3D Poses for recognizing ADL? To this end, we propose an\nextension of a pose driven attention mechanism: Video-Pose Network (VPN),\nexploring two distinct directions. One is to transfer the Pose knowledge into\nRGB through a feature-level distillation and the other towards mimicking pose\ndriven attention through an attention-level distillation. Finally, these two\napproaches are integrated into a single model, we call VPN++. We show that\nVPN++ is not only effective but also provides a high speed up and high\nresilience to noisy Poses. VPN++, with or without 3D Poses, outperforms the\nrepresentative baselines on 4 public datasets. Code is available at\nhttps://github.com/srijandas07/vpnplusplus.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC