Research on Passenger Posture Prediction in Subway Station Based on Lightweight Improved YOLOv8-Pose
To address challenges including severe occlusion, large-scale variation, inaccurate keypoint localization, and strict real-time constraints in subway surveillance scenarios, this paper proposes a lightweight human pose estimation framework based on an enhanced YOLOv8-Pose architecture. The proposed method introduces four key components: a DWR (Dilation-wise Residual) module to expand multi-scale receptive fields for improved joint context modeling, an HFERB (High-Frequency Enhancement Residual Block) to recover fine-grained structural details lost during downsampling, an OrthoNets-based orthogonal attention mechanism to reduce channel redundancy while preserving discriminative representations, and an MFEB (Multi-Scale Feature Enhancement Block) to enable adaptive feature fusion under varying crowd densities. Extensive experiments on the COCO 2017 dataset demonstrate that the proposed method achieves 87.9% mAP@0.5 and 64.9% mAP@0.5–0.95, improving the baseline YOLOv8-Pose by 1.7% and 1.0%, respectively, while reducing validation pose loss by 5.77%. After lightweight optimization with ONNX export and FP16 inference, the model is compressed to 3.99M parameters and achieves 76 FPS inference speed on the same hardware, significantly improving efficiency compared with the baseline, while achieving 38 FPS on the full model. Both configurations satisfy real-time constraints for subway surveillance deployment. Qualitative results on real subway surveillance videos further demonstrate strong robustness under occlusion and scale variation. Overall, the proposed method achieves a favorable trade-off between accuracy and efficiency, making it suitable for real-time human pose estimation in intelligent transportation systems.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex