CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition

Skeleton-based multi-entity action recognition is a challenging task aiming to identify interactive actions or group activities involving multiple diverse entities. Existing models for individuals often fall short in this task due to the inherent distribution discrepancies among entity skeletons, leading to suboptimal backbone optimization. To this end, we introduce a Convex Hull Adaptive Shift based multi-Entity action recognition method (CHASE), which mitigates inter-entity distribution gaps and unbiases subsequent backbones. Specifically, CHASE comprises a learnable parameterized network and an auxiliary objective. The parameterized network achieves plausible, sample-adaptive repositioning of skeleton sequences through two key components. First, the Implicit Convex Hull Constrained Adaptive Shift ensures that the new origin of the coordinate system is within the skeleton convex hull. Second, the Coefficient Learning Block provides a lightweight parameterization of the mapping from skeleton sequences to their specific coefficients in convex combinations. Moreover, to guide the optimization of this network for discrepancy minimization, we propose the Mini-batch Pair-wise Maximum Mean Discrepancy as the additional objective. CHASE operates as a sample-adaptive normalization method to mitigate inter-entity distribution discrepancies, thereby reducing data bias and improving the subsequent classifier's multi-entity action recognition performance. Extensive experiments on six datasets, including NTU Mutual 11/26, H2O, Assembly101, Collective Activity and Volleyball, consistently verify our approach by seamlessly adapting to single-entity backbones and boosting their performance in multi-entity scenarios. Our code is publicly available at https://github.com/Necolizer/CHASE .

Paper

Similar papers

Peer review

Reviewer WotV6/10 · confidence 4/52024-07-09

Summary

This paper tackled the issue of inter-entity distribution discrepancies in multi-entity action recognition. The authors proposed convex hull adaptive shift method to minimize the cross entity discrepancies, where CLB and MPMMD are proposed to assist the learning procedure. The method is verified to be effective among various datasets and backbones.

Strengths

1.This paper proposed an interesting idea by using implicit convex hull as constraints to achieve adaptive coordinate shift. 2.The proposed approach is verified to be effective on various datasets and backbones. 3.This method can serve as a good contribution to the skeleton-based human action recognition community.

Weaknesses

1. The introduction section should be improved. For example on line 54, why do we need to achieve the discrepancy minimization? on line 52, why do we need to achieve sample adaptive coefficients? The motivation should be highlighted. The links among these proposed items should be also improved on line 57-60 and need more insights. 2. The novelty of the CLB is limited. The format of attributes learning shown in Eq.8 is commonly used to construct concept learners. What is the difference between the concept bottleneck [1] and the CLB? If you use the concept learner from some existing works, e.g., [2], will it be better than CLB? [1] Shin S, Jo Y, Ahn S, et al. A closer look at the intervention procedure of concept bottleneck models[C]//International Conference on Machine Learning. PMLR, 2023: 31504-31520. [2] Wang B, Li L, Nakashima Y, et al. Learning bottleneck concepts in image classification[C]//Proceedings of the ieee/cvf conference on computer vision and pattern recognition. 2023: 10962-10971. 3. More insights should be given in Section 4.2. Why does the proposed method help? The authors are encouraged to enrich the analysis. The authors are encouraged to discuss the computational complexity brought by the proposed method. 4. The authors are encouraged to discuss the computational complexity brought by the proposed method.

Questions

1. Improving the Introduction Section: a. Why is it necessary to achieve discrepancy minimization (line 54)? b.Why do we need to achieve sample adaptive coefficients (line 52)? c. Can the authors highlight the motivation behind these needs and improve the links among the proposed items (lines 57-60) with more insights? 2. Novelty of the CLB: a. How does the concept bottleneck (CLB) differ from the commonly used format of attributes learning in concept learners, such as in Eq. 8? b. What are the differences between the CLB and the concept bottleneck models discussed in Shin et al. (2023)? c. If the concept learner from existing works (e.g., Wang et al. (2023)) were used, would it perform better than the CLB? 3. Insights in Section 4.2: a. Why does the proposed method provide benefits? b. Can the authors provide a more detailed analysis to enrich Section 4.2? 4. Computational Complexity: a. What is the computational complexity introduced by the proposed method?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

yes in supplementary

Authorsrebuttal2024-08-12

Looking Forward to Further Discussions

Dear Reviewer, Thanks again for your insightful comments on our paper. We have submitted the response to your comments and the global response with a PDF file. Please let us know if you have additional questions so that we can address them during the discussion period. We hope that you can consider rasing the score. Thank you

Reviewer WotV2024-08-12

To the authors

Dear authors, thank you very much for your rebuttal. I think most of my concerns are handled and I will improve my score to 6. Best,

Authorsrebuttal2024-08-13

Thank You for Your Positive Feedback and Consideration

Dear Reviewer, Thank you for your positive feedback and for taking the time to review our rebuttal. We're glad that our responses addressed your concerns, and we appreciate your willingness to improve the score. Thank you again for your thoughtful consideration. Best regards,

Reviewer 16iH6/10 · confidence 4/52024-07-10

Summary

This paper proposes CHASE, a multi-entity skeleton data augmentation/preprocessing technique, to mitigate inter-entity distribution gaps and improve the multi-entity action recognition. Specifically, the authors formulate a new constraint called ICHAS, design a lightweight block CLB to learn the nonlinear mapping from input to the weight matrix in ICHAS, and introduce an objective to guide the discrepancy minimization in CLB training. The authors conduct comprehensive experiments to show the effectiveness of the method.

Strengths

a) The paper is well-written, easy to understand, and well-organized. b) The method is well-motivated, well-ablated, and the experiments are presented clearly.

Weaknesses

a) There’re some typos need to be revised carefully, e.g. l131 “be be”, Table 4 “-68.56”. b) Regarding clarity, the captions of Figures and Tables can be improved.

Questions

a) Figure 2 is very informative and good for understanding the whole method. However, there are still some notations needed to be explained, e.g. the green and red circles under the two skeleton coordinates. It would be helpful if the authors introduce the method details (in sections 3.1 to 3.3) with references to specific parts of Figure 2. b) It’s good to have multiple runs and report the standard deviation. How many seed initializations exactly do the authors use? c) From Table 1, the performance margin varies a lot. On H2O and CAD, the top-1 accuracies have increased by over 9%. Could the authors explain or provide some discussion on this point? d) From Table 7 in the appendix, it seems like the authors use 25 ground-truth 3D joints as skeleton inputs for NTU-60 and NTU-120 datasets. Since 2D estimated joints tend to achieve better action recognition performance in dominant models, do the authors verify CHASE using 17 estimated 2D joints with COCO layout as input? e) Table 6 shows the mixed recognition results on the entire NTU-120 dataset with X-Sub setting. Table 8 in the appendix shows the mixed recognition results on NTU-120 X-Sub and X-Set settings. Why not replace Table 6 with Table 8? Similar to Table 5 and Table 9, why not replace Table 5 with Table 9 or integrate the number of parameters into Table 1 (may discard the column of ‘Venue’)? Small details (no need to address them, just suggestions): a) Regarding equations (9) and (10), it would be better to clarify the definitions of sup(*) and C(E,2).

Rating

6

Confidence

4

Soundness

2

Presentation

3

Contribution

2

Limitations

Although the method has been verified on six benchmarks, these datasets are relatively small, with the number of categories ranging from 4 to 36, except Assembly 101 has 1380 action categories. However, the performance gain on Assembly 101 is very small (<=0.21%). The reviewer has some concerns about the generalization of this method.

Authorsrebuttal2024-08-12

Looking Forward to Further Discussions

Dear Reviewer, Thanks again for your insightful comments on our paper. We have submitted the response to your comments and the global response with a PDF file. Please let us know if you have additional questions so that we can address them during the discussion period. We hope that you can consider rasing the score. Thank you

Reviewer hFtj4/10 · confidence 4/52024-07-12

Summary

This paper focuses on the interesting problem of the normalization strategy for multi-entity skeletons in skeleton-based action recognition. The proposed method is intuitive, and the authors provided detailed implementation details for reproduction. However, this work is unclear, and the experiments are unconvincing.

Strengths

This paper focuses on the interesting problem of the normalization strategy for multi-entity skeletons in skeleton-based action recognition. The proposed method is intuitive, and the authors provided detailed implementation details for reproduction.

Weaknesses

(1) The purpose of the multi-entity action recognition task is not clear. Is it to recognize each individual’s action or to classify group activities? If the purpose varies across different datasets, please clarify this in the experiment. Additionally, I am curious whether the optimal normalization strategy differs for these two purposes. For example, the method used in S2CoM seems more suitable for recognizing each individual’s action. (2) What is the main difference between the proposed method and the simple strategy of shifting the origin of multi-entities to their common center? Please add a comparison experiment with this method. (3) Although the normalization strategy is an important trick and can bring significant improvement in the action classification task, I still have a concern about whether it is worth conducting an additional network to achieve this simple trick by introducing extra learnable parameters. Some heuristic strategies may be more efficient and general. Accordingly, can the proposed module be transferred among different datasets without retraining the module?

Questions

What is the main difference between the proposed method and the simple strategy of shifting the origin of multi-entities to their common center? Please add a comparison experiment with this method.

Rating

4

Confidence

4

Soundness

2

Presentation

2

Contribution

2

Limitations

(1) The purpose of the multi-entity action recognition task is not clear. Is it to recognize each individual’s action or to classify group activities? If the purpose varies across different datasets, please clarify this in the experiment. Additionally, I am curious whether the optimal normalization strategy differs for these two purposes. For example, the method used in S2CoM seems more suitable for recognizing each individual’s action. (2) What is the main difference between the proposed method and the simple strategy of shifting the origin of multi-entities to their common center? Please add a comparison experiment with this method. (3) Although the normalization strategy is an important trick and can bring significant improvement in the action classification task, I still have a concern about whether it is worth conducting an additional network to achieve this simple trick by introducing extra learnable parameters. Some heuristic strategies may be more efficient and general. Accordingly, can the proposed module be transferred among different datasets without retraining the module?

Authorsrebuttal2024-08-12

Looking Forward to Further Discussions

Dear Reviewer, Thanks again for your insightful comments on our paper. We have submitted the response to your comments and the global response with a PDF file. Please let us know if you have additional questions so that we can address them during the discussion period. We hope that you can consider rasing the score. Thank you

Authorsrebuttal2024-08-13

Clarifications on Transferability and Multi-Entity Action Recognition

Dear Reviewer, Thank you for your feedback. We apologize for not fully addressing your concerns in our initial response. 1. We have followed your suggestions, conducting the experiments to evaluate the transferability of our CHASE module across different datasets without retraining. Specifically, we first trained the CHASE + CTR-GCN backbone on the challenging ASB101 dataset, achieving a top-1 accuracy of 28.03%. **We then transferred the CHASE module to the H2O (two-hand version) dataset without retraining, where it achieved 56.61% accuracy.** **This result outperforms both retraining the module on H2O (56.47%) and training only the CTR-GCN backbone on H2O (48.48%).** These findings demonstrate that our proposed module can indeed be transferred among different datasets without retraining. 2. Multi-entity action recognition **aims to classify interactions involving multiple entities, which could be people, objects, or other elements within a scene.** Examples of such actions include *cheers and drink*, *exchanging things*, *walking apart,* and *talking*. Group activities are a subset of multi-entity actions [73, 75, 76]. Unlike traditional action recognition, which typically focuses on a single subject performing a single action, multi-entity action recognition addresses the complexity of interpreting actions that involve multiple participants or objects interacting simultaneously. We hope this clarification addresses the confusion. We appreciate your insights and hope this additional information helps to resolve your concerns. Best regards,

Reviewer N2tt6/10 · confidence 3/52024-07-12

Summary

The paper proposes a normalization method for skeleton-based multi-entity recognition based on finding the center of mass within the convex hull of the spatio-temporal domain of the point cloud defined by the skeletons over a sequence. The main idea is to "center" the world of skeletons to unbias the subsequent detector and boost their performance. Building on this motivation, the authors find that a fixed, learnable parameterized network can be used for that purpose, facilitating the inference. The experiments demonstrate that their proposed algorithm boosts performance over the corresponding baselines

Strengths

The paper is technically sound and the authors follow a proper mathematical derivation that leads to the design of CHASE in a clever manner. The paper includes an extensive supplementary material with code and further analysis that make the paper rather complete. The method is elegant and simple, providing with a very efficient, lightweight network that shows provable performance on a broad variety of datasets. Ablation studies are conducted to validate the proposed parts, as well as to compare against other normalization alternatives. The paper is well documented with an extensive coverage of related literature.

Weaknesses

Overall I believe that the presentation should be improved, clearly stating the contribution and motivation for it. It takes a good read to understand that what the authors are proposing as a lightweight network is the result of mathematically deriving an iterative approach for normalization of skeletons within their convex hull. It should be clearly stated that the method aims to accompany other methods for multi-entity activity recognition by adding an extra normalization step, which consists of what is presented in Section 3. Similarly, the sketch depicted in Figure 1 is a bit confusing and does not really illustrate what the authors aim to solve. I would suggest the authors to provide a clear motivation example that leads to their method. The caption in Fig. 1 is rather poorly written (I did not understand it at least).

Questions

I would insist on the authors to please elaborate a bit better on the motivation and an example where former normalization leads to the classifiers to produce wrong results, with their method mitigating such problem.

Rating

6

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

N/A

Authorsrebuttal2024-08-12

Looking Forward to Further Discussions

Dear Reviewer, Thanks again for your insightful comments on our paper. We have submitted the response to your comments and the global response with a PDF file. Please let us know if you have additional questions so that we can address them during the discussion period. We hope that you can consider rasing the score. Thank you

Reviewer N2tt2024-08-13

Answer

I am happy with the provided response and I have no further questions in regards to it. I believe this paper lies above the acceptance threshold.

Reviewer hFtj2024-08-13

Rebuttal Comment

I keep my initial score, where the authors did not struggle to handle my concerns, like "transferred among different datasets without retraining the module". Besides, the multi-entity action recognition task is still confusing.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC