EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
There is an increasing demand for lightweight multi-person pose estimation\nfor many emerging smart IoT applications. However, the existing algorithms tend\nto have large model sizes and intense computational requirements, making them\nill-suited for real-time applications and deployment on resource-constrained\nhardware. Lightweight and real-time approaches are exceedingly rare and come at\nthe cost of inferior accuracy. In this paper, we present EfficientHRNet, a\nfamily of lightweight multi-person human pose estimators that are able to\nperform in real-time on resource-constrained devices. By unifying recent\nadvances in model scaling with high-resolution feature representations,\nEfficientHRNet creates highly accurate models while reducing computation enough\nto achieve real-time performance. The largest model is able to come within 4.4%\naccuracy of the current state-of-the-art, while having 1/3 the model size and\n1/6 the computation, achieving 23 FPS on Nvidia Jetson Xavier. Compared to the\ntop real-time approach, EfficientHRNet increases accuracy by 22% while\nachieving similar FPS with 1/3 the power. At every level, EfficientHRNet proves\nto be more computationally efficient than other bottom-up 2D human pose\nestimation approaches, while achieving highly competitive accuracy.\n
Paper
References (73)
Scroll for more · 38 remaining