The Center of Attention: Center-Keypoint Grouping via Attention for Multi-Person Pose Estimation

We introduce CenterGroup, an attention-based framework to estimate human\nposes from a set of identity-agnostic keypoints and person center predictions\nin an image. Our approach uses a transformer to obtain context-aware embeddings\nfor all detected keypoints and centers and then applies multi-head attention to\ndirectly group joints into their corresponding person centers. While most\nbottom-up methods rely on non-learnable clustering at inference, CenterGroup\nuses a fully differentiable attention mechanism that we train end-to-end\ntogether with our keypoint detector. As a result, our method obtains\nstate-of-the-art performance with up to 2.5x faster inference time than\ncompeting bottom-up methods. Our code is available at\nhttps://github.com/dvl-tum/center-group .\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC