Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, existing Multimodal Large Language Models (MLLMs) face challenges in integrating audio and recognizing subtle facial micro-expressions. To address this, we introduce the MERR dataset, containing 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories. This dataset enables models to learn from varied scenarios and generalize to real-world applications. Furthermore, we propose Emotion-LLaMA, a model that seamlessly integrates audio, visual, and textual inputs through emotion-specific encoders. By aligning features into a shared space and employing a modified LLaMA model with instruction tuning, Emotion-LLaMA significantly enhances both emotional recognition and reasoning capabilities. Extensive evaluations show Emotion-LLaMA outperforms other MLLMs, achieving top scores in Clue Overlap (7.83) and Label Overlap (6.25) on EMER, an F1 score of 0.9036 on MER2023-SEMI challenge, and the highest UAR (45.59) and WAR (59.37) in zero-shot evaluations on DFEW dataset.

Paper

Similar papers

Peer review

Reviewer 646t5/10 · confidence 4/52024-06-22

Summary

This paper implements multimodal emotion recognition and reasoning by fine-tuning the LLaMA model with instructions. It is trained on a large-scale dataset, fine-tuned, and tested on three datasets.

Strengths

The global, temporal and local features of the video modality are considered, and LoRA fine-tuning (all tokens), Prompt fine-tuning (text modality) and supervised fine-tuning (linear layer) are used at the same time.

Weaknesses

The innovation is limited and only existing methods are used. The key problems are not solved.

Questions

1. The complex interaction relationship between modalities is not considered. 2. The process is complicated, and the computing resource requirements are high. 3. The effect of Prompt Tuning and Instruction Tuning is highly dependent on the design of prompt words and instructions. If the prompt words or instructions are not accurate or comprehensive enough, it may impact the performance of the model. 4. The lack of reproducibility: the anonymous Github link provided by the author is empty.

Rating

5

Confidence

4

Soundness

2

Presentation

2

Contribution

1

Limitations

Yes

Authorsrebuttal2024-08-10

Response to Reviewer 646t

Dear Reviewer, Thank you for your insights. We have made the following revisions to address your concerns: 1. We clarified the interaction between modalities in our approach, demonstrating the effectiveness of our method. 2. We discussed the computational efficiency of our approach and how we ensured that the process is resource-efficient. 3. We ensured the reproducibility of our work by verifying that all necessary files are accessible in our anonymous repository. 4. We highlighted the innovative aspects of our methodology, addressing the challenges in multimodal emotion recognition. We hope these updates address your concerns. We appreciate any further feedback you might have. Best regards, The Authors

Authorsrebuttal2024-08-12

Follow-Up on Revisions and Inquiry on Additional Concerns

Dear Reviewer 646t, Thank you for your thorough review and for raising the score, which we greatly appreciate. As we approach the rebuttal deadline, we wanted to check if there are any remaining concerns or questions that we could address. Your feedback has been invaluable, and we are committed to making any further necessary improvements. Please let us know if there is anything else we should consider. Best regards, The Authors

Authorsrebuttal2024-08-13

Follow-Up on Revisions and Interaction

Dear Reviewer 646t, Thank you for your thorough review and for raising the score, which we greatly appreciate. We have invested a significant amount of effort into this work, and your feedback has been instrumental in guiding our revisions. As we approach the rebuttal deadline, we wanted to ensure that all your concerns have been adequately addressed. We would also like to invite you to try out our demo, available in the anonymous repository. Our work has already attracted considerable attention, leading to a high number of visits to the demo, which has significantly increased the maintenance costs. Despite this, we have kept it running during the review process to provide valuable insights into the practical application of our methods. If there are any remaining questions or additional feedback you could provide, we would be more than happy to address them. Best regards, The Authors

Reviewer V8m95/10 · confidence 3/52024-07-10

Summary

The paper introduces a multimodal large language model, named Emotion-LLaMA for emotional state understanding. The authors use open-source tools to collect and annotate a dataset, named MERR for model pre-training. Then they perform instruction-tuning on downstream datasets for emotion recognition and emotion reasoning. Extensive experiments are conducted and demonstrate the promising performance of the proposed approach.

Strengths

1. Extensive experiments are conducted and the method achieves SOTA performance on various dataset for emotion recognition and reasoning. 2. Straightforward visualizations are presented.

Weaknesses

1. The paper claims the MERR dataset as a core contribution. However, there is no systematic evaluation for the label quality of the dataset. I understand that pre-training on this dataset improves the downstream performance and thus its validity can be shown to some extent. However, this may be due to the diversity of the unlabeled data instead of the automatically generated label. 2. Table 3, with audio and video inputs, Emotion-LLaMA’s performance is worse/close to the baseline (VAT).

Questions

1. How do you pre-train Emotion-LLaMA? Is it supervised learning with coarse-grained emotion labels? 2. Do you plan to release the MERR dataset? 3. Table 2 fine-tuning, why does Emotion-LLaMA get worse performance on Disgust than MAE-DFER and VideoMAE?

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Please refer to weaknesses

Authorsrebuttal2024-08-10

Response to Reviewer V8m9

Dear Reviewer, Thank you for your thoughtful review. We have made the following revisions to address your feedback: 1. We provided a detailed breakdown of our pre-training and fine-tuning processes. 2. We discussed the impact of the MERR dataset on the model’s performance and the challenges associated with recognizing the 'disgust' emotion category. 3. We ensured that the MERR dataset is now fully accessible and clarified its role in our methodology. We hope these revisions address your concerns. Please let us know if further clarification is needed. Best regards, The Authors

Authorsrebuttal2024-08-12

Inquiry on Additional Concerns

Dear Reviewer V8m9, Thank you for your thorough review and constructive feedback. We appreciate the time you have taken to assess our work and provide insights that have greatly contributed to improving our paper. As the rebuttal period comes to a close, we wanted to ensure that all your concerns have been adequately addressed. If there are any remaining questions or points that need further clarification, please let us know, and we will be happy to provide additional information. Your feedback is invaluable to us, and we are committed to making the necessary improvements to our submission. Best regards, The Authors

Reviewer V8m92024-08-13

Dear authors, Thank you for your hard work, and I apologize for the delayed response. My major concerns have been addressed, and I will maintain my ratings and vote to accept the paper.

Authorsrebuttal2024-08-13

Consideration Request

Dear Reviewer V8m9, Thank you for your thoughtful review and constructive feedback. We appreciate the time you have taken to assess our work and provide insights that have greatly contributed to improving our paper. We have thoroughly addressed all the issues you raised, and we believe these revisions have significantly enhanced our manuscript. Given that all your concerns have been carefully considered and rectified, we respectfully inquire if you could consider raising the score. Additionally, we encourage you to explore the demo available through our anonymous submission. Although maintaining the demo incurs significant daily costs—especially now that our work has gained some traction—we have decided to keep it running during the review process to ensure you and other reviewers have full access to it. If you have any remaining questions or additional suggestions that could further improve our submission, we would be grateful for your feedback. Once again, we would like to express our gratitude for your commitment to reviewing our paper and for the constructive comments that have guided our revisions. Best regards, The Authors

Reviewer iNoC3/10 · confidence 4/52024-07-12

Summary

The paper presents a new multi-modal instruction tuning dataset for emotion recognition. The authors also present results of training on this dataset with a multi-modal architecture based on LLaMa 2. They show evaluation results on DFEW and MER2023.

Strengths

Evaluating multi-modal emotion recognition approaches based on LLMs are a very relevant research topic. The results shown by the authors are convincing w.r.t. the performance on emotion recognition datasets.

Weaknesses

------------- The argument presented in lines 31 and 33 is not convincing. Why does an inability in methods lead to a lack of datasets? Related Work: ------------- As the dataset is one of the claimed contributions of the paper, there should be a discussion on previous emotion recognition datasets. It is important to lay out the reasons why these previous datasets cannot be used (or not easily be used) for instruction tuning of LLMs. Methodology: ------------ What videos is the MERR dataset based on? I could not find details in the paper on the selection process. Judging from the screenshots, it seems to be movie clips. This has important implications on the concept of "emotion" that is addressed in the paper and needs to be clarified. The chosen methodology lacks justification. For example, what is the reasoning behind selecting the frames with the highest sum of AU activations across all AUs included in OpenFace? It seems to me that there would be the danger of having a strong bias towards moments when the person is speaking, as this often leads to high AU activations, especially AUs related to the lower half of the face. The mapping of AUs to facial expression labels needs more explanation. E.g. is "happy" assigned if the combination of AU06, AU12, AU14 is active, or is it also assigned if only some of those AUs are active? Figure 1 does not show a facial expression description according to Table 7 at all, only descriptions according to Table 8. Concerning Table 8, there are several descriptions given for each AU. Are they all used at the same time, or is only one of them chosen? In Figure 1 it appears that only one of them is chosen, but how is this decided? Several further steps are not clearly defined, e.g. how does LLaMA-3 "refine" the annotations by aggregating the multimodal description? How is the instruction following data constructed of which a single example is presented in Table 9. Is it done manually? If yes, how and how many samples are created? In 3.1 it is not clear, how the dataset is "auto-annotated" with emotion labels. It is also not clear how these annotations are refined (with human involvement)? In general the concept of "emotion" used in the paper remains unclear. Is it about emotional displays, about internal states,...? In later parts of the method section it seems authors are targeting (internal?) emotional states. It would be important to know how they were annotated. Do the instructions for multimodal emotion recognition (Table 11) refer to different tasks? Some of these instructions seem to target displayed emotions, some target internal states. Training details are unclear. With the description provided in the paper it is difficult to understand the training procedure. Evaluation: ----------- The authors employ chatGPT to evaluate emotion reasoning. It is not clear to what extent this approach leads to a valid evaluation. E.g. ChatGPT was shown to be biased on emotion-related tasks [1]. In the end, evaluating the emotion understanding capabilities of one language model using another language model is circular. The impact of using the proposed MERR dataset for pre-training needs to be evaluated. [1] R. Mao, Q. Liu, K. He, W. Li, and E. Cambria, “The biases of pre-trained language models: An empirical study on prompt-based sentiment analysis and emotion detection,” IEEE Transactions on Affective Computing, 2022.

Questions

-

Rating

3

Confidence

4

Soundness

2

Presentation

2

Contribution

2

Limitations

Limitations are not discussed in enough detail. There is no separate limitations section. The authors also do not lay out the limitations concerning e.g. (citing from the checklist): "The authors should reflect on the scope of the claims made, e.g., if the approach was580 only tested on a few datasets or with a few runs. In general, empirical results often581 depend on implicit assumptions, which should be articulated." "The authors should reflect on the factors that influence the performance of the approach.583 For example, a facial recognition algorithm may perform poorly when image resolution584 is low or images are taken in low lighting. Or a speech-to-text system might not be585 used reliably to provide closed captions for online lectures because it fails to handle586 technical jargon."

Authorsrebuttal2024-08-10

Response to Reviewer iNoC

Dear Reviewer, Thank you for your constructive comments. We have revised the manuscript to address your concerns: 1. We included a detailed discussion on previous emotion recognition datasets and clarified why they are not suitable for instruction tuning of large language models. 2. We provided a clearer explanation of our methodology, including the MERR dataset’s annotation process and its role in enhancing our model. 3. We added more details about our training process and provided a rationale for using ChatGPT in the evaluation. We hope these changes address your concerns. We welcome any additional feedback you may have. Best regards, The Authors

Authorsrebuttal2024-08-12

Further Clarifications and Inquiry on Remaining Concerns

Dear Reviewer iNoC, Thank you for your detailed review and the constructive feedback you provided. We have carefully addressed your concerns in the revised manuscript, including a comprehensive discussion of previous emotion recognition datasets, a clearer explanation of our methodology, and the rationale behind our use of ChatGPT for evaluation. We understand that your score reflects concerns, and we deeply appreciate your critical assessment. To ensure we have addressed all your points, we would like to know if there are any remaining misunderstandings or unclear aspects that we can further clarify before the rebuttal deadline. Your insights have been invaluable in improving our work, and we are committed to making any necessary adjustments. Best regards, The Authors

Authorsrebuttal2024-08-13

Further Clarifications

Dear Reviewer iNoC, Thank you for your detailed review and constructive feedback. We've carefully addressed your concerns in the revised manuscript, including a thorough discussion of previous datasets, a clearer explanation of our methodology, and the rationale for using ChatGPT in evaluation. We understand your score reflects some concerns, and we believe there may be some misunderstandings about our work. We have put significant effort into this project and are eager to clarify any remaining points. We invite you to explore our demo in the anonymous repository and are happy to address any further questions or concerns you might have. Best regards, The Authors

Reviewer iNoC2024-08-14

No change in evaluation

I read the rebuttal and will remain with my score. Justification below. A2: DialogueLLM uses MELD, IEMOCAP and EMoryNLP. All these three datasets are not mentioned in the author's response but have been used for instruction tuning LLMs. A3: The selection procedure is still unclear - i.e. how was the dataset sampled from MER2023? A4 & A5: No references are given for the supposed connection between AUs and emotion expression. It is unclear how stable this connection is in different contexts. A6 & A7: The author's say "Llama-3 refines annotations" but how this is exactly done and how the quality of this refinement can be assured is unclear. What is the protocol for human annotators. What are human "preferences" here? A9: Training details that are needed to understand the approach should be part of the main paper. A10: There should be a proper human evaluation of this process, at least on a part of the dataset. The issue about internal emotional states that I raised is not mentioned in the rebuttal.

Reviewer 3ouG4/10 · confidence 3/52024-07-13

Summary

The paper presents the Emotion-LLaMA model, a multimodal emotion recognition and reasoning system that integrates audio, visual, and textual inputs. - The authors constructed the MERR dataset, which includes 28,618 coarse-grained and 4,487 fine-grained annotated samples across diverse emotional categories, enabling models to learn from varied scenarios and generalize to real-world applications. - The Emotion-LLaMA model incorporates specialized encoders for audio, visual, and textual inputs, aligning the features into a modified LLaMA language model and employing instruction tuning to enhance both emotional recognition and reasoning capabilities. - Extensive evaluations show that Emotion-LLaMA outperforms other multimodal large language models, achieving top scores on the EMER, MER2023, and DFEW datasets. The main contributions are: - The MERR dataset, a valuable resource for advancing large-scale multimodal emotion model training and evaluation. - The Emotion-LLaMA model, which excels in multimodal emotion recognition and reasoning through the innovative use of instruction tuning. - Establishing Emotion-LLaMA as the current state-of-the-art model in public competitions for multimodal emotion analysis.

Strengths

- The paper is well structured and well presented. The author organizes the article into five sections: introduction, related work, methodology, experiments, and conclusion. They clearly describe how they conduct data annotation and model design, provide a good introduction to the experimental setup and analysis of the experimental results, and present a relatively clear conclusion. - Model details are thorough. The authors have provided a good description of their model and training methods, allowing me to clearly understand how the model is designed and trained. I believe their results are reproducible. - This work is valuable . The authors provided a paradigm for emotion annotation of multimodal data and an annotated dataset. They also offered a clear explanation of the annotation process, which will contribute to the development of the related field.

Weaknesses

- Insufficient experiments: Although the author has conducted some comparative and ablation experiments, it is obviously insufficient for such a complex multimodal LLM. - Missing details of experimental setup: The authors mentioned that they fine-tuned on several target datasets, but the details of the fine-tuning (including data volume, dataset division, and fine-tuning setup) were not included in the article. This can lead to a decrease in the credibility of their results. - Lack of result analysis: Although the proposed model surpasses existing models in many metrics, the authors only list their experimental results in the results section without further analysis and explanation of the results and some phenomena. The lack of proper explanation and analysis can make some results confusing.

Questions

1. In section 4.2, the author mentioned that they used the HuBERT-Chinese large model for audio modality input processing. Regarding this part, I would like to know the following questions: - Are the experimental results sensitive to language? I hope the author can provide more experimental results to illustrate this point. - In the ablation study section (Tab5), the author seems to have only conducted ablation on the visual encoder. Why didn't they further conduct ablation experiments on the audio encoder, considering that there are many alternatives to Hubert? - Based on the previous question, does the author believe that the audio modality is not important in this task? I would like to see more ablation results on the modality scale. 2. As I mentioned in the weaknesses part, could the author provide more details of fine-tuning on the target datasets? 3. Is Multimodal Emotion Recognition a classification task? How do the authors explain the Dis column in Table 2?

Rating

4

Confidence

3

Soundness

3

Presentation

3

Contribution

4

Limitations

Although the author mentioned in the checklist that they discussed the limitations of the article, they did not explicitly discuss them in the text. Meanwhile, I noticed that the author did not mention their data sources in the article, and I am concerned whether this might involve data copyright issues.

Authorsrebuttal2024-08-10

Response to Reviewer 3ouG

Dear Reviewer, Thank you for your valuable feedback. We have carefully considered your comments and made the following revisions to address your concerns: 1. We added additional ablation studies on the audio modality to provide more comprehensive insights. 2. We clarified the fine-tuning process and included more details to improve transparency. 3. We expanded our analysis of the results, particularly focusing on the challenges and future improvements regarding the 'disgust' emotion category. We hope these revisions meet your expectations. Please let us know if there are any further issues or if additional clarification is needed. Best regards, The Authors

Authorsrebuttal2024-08-12

Follow-Up on Revisions and Inquiry on Remaining Concerns

Dear Reviewer 3ouG, Thank you for your valuable feedback and for taking the time to carefully review our submission. We have thoughtfully considered your comments and made several revisions to address the issues you raised, including additional ablation studies on the audio modality, clarifying the fine-tuning process, and expanding our analysis of the results. As we approach the rebuttal deadline, we would like to ensure that all your concerns have been adequately addressed. If there are any remaining issues or areas where you believe further clarification is needed, please let us know. We are committed to making any necessary improvements to our work. We appreciate your efforts in helping us enhance the quality of our paper. Best regards, The Authors

Authorsrebuttal2024-08-13

Follow-Up on Revisions and Demo Interaction

Dear Reviewer 3ouG, Thank you once again for your valuable feedback and for taking the time to carefully review our submission. We have made several revisions to address the issues you raised, including additional ablation studies on the audio modality, clarifying the fine-tuning process, and expanding our analysis of the results. As we approach the final stages of the rebuttal process, we wanted to ensure that all your concerns have been fully addressed. We also invite you to explore our demo, available in the anonymous repository. Our work has already gained some traction, leading to a high number of visits to the demo, which has significantly increased the maintenance costs. Despite this, we have kept it running to provide full access during the review process. We believe it could offer further insights into our work, and we would greatly appreciate any feedback you might have. We are eager to interact with you further and are committed to making any additional improvements necessary. Best regards, The Authors

Authorsrebuttal2024-08-14

Response to Ethics Review

Thank you for your detailed and insightful review. Below, we address each of your concerns and outline the actions we have taken or plan to take in the revised manuscript. **Data Privacy and Consent:** The MERR dataset is sourced from the MER2023 dataset, consisting of over 70,000 unannotated samples from publicly available movies and TV series. We have signed the necessary End User License Agreements (EULA) and received permission from the original data providers. Since this data originates from publicly available media, and participant consent was obtained during the original production, additional consent for our research is not required. We will also contact the original dataset's authors to ensure compliance with ethical standards. A section will be added to the appendix detailing the video data licensing and participant consent context. **Fairness and Bias:** The MERR dataset is designed to cover a wide range of multimodal emotional cues while avoiding discriminatory content related to sensitive attributes. We acknowledge the potential for inherent biases and have implemented measures to assess and mitigate them. We provided detailed scoring criteria and results in Tables 14 and 15, confirming the absence of significant biases. We used ChatGPT as a neutral evaluator, focusing on textual annotations to reduce bias related to visual or demographic factors. Moving forward, we will explore additional validation methods to ensure fairness. **Reproducibility and Transparency:** The open-source MERR dataset includes automatically generated emotion description JSON files, excluding original video content due to licensing restrictions. Researchers can access the original video data by applying directly to the MER2023 providers and signing the EULA. The revised manuscript will clearly outline the terms of use and provide detailed reproduction instructions. Our GitHub repository now includes all necessary scripts, documentation, and a README file to guide the reproduction process, with access restricted to research purposes only. **Broader Impact:** We have reviewed the MERR dataset and Emotion-LLaMA outputs and found no evident discrimination or bias. However, we recognize the potential risks of reinforcing societal biases in real-world applications. We will investigate these impacts further and discuss them explicitly in the revised manuscript. Future work will incorporate advanced bias detection techniques and input from ethics and social sciences experts. In summary, we believe the steps outlined address the ethical concerns raised. We are committed to ensuring our research adheres to ethical standards and will incorporate these considerations into the revised manuscript. Thank you for your thoughtful feedback.

Authorsrebuttal2024-08-17

Follow-Up on Ethics Review

Dear Ethics Reviewer gVR7, We have carefully addressed each of your concerns in our previous response and have outlined the steps we are taking to incorporate these considerations into the revised manuscript. We wanted to follow up to ensure that we have fully addressed your concerns. If there are any additional issues or points of clarification needed, please let us know. We are committed to adhering to the highest ethical standards in our research and are happy to provide further information if required. We appreciate your time and input in this review process. Sincerely, Authors

Program Chairs2024-08-14

Hi all, the author-reviewer discussion for this paper is extended to Aug 16 11:59pm ET for authors to reply to ethics reviews. -- PCs

Authorsrebuttal2024-08-17

Detailed Response to Reviewer iNoC's Feedback (Part-1)

Dear Reviewer iNoC, Thank you for your detailed feedback and for engaging with our work on Emotion-LLaMA. We acknowledge that your additional questions and concerns were raised just a few hours before the rebuttal deadline. Unfortunately, due to these time constraints, we were unable to provide a comprehensive response at that time. We greatly appreciate the opportunity to now address each of your points thoroughly. --- **Q2: DialogueLLM uses MELD, IEMOCAP, and EmoryNLP. All these three datasets are not mentioned in the author's response but have been used for instruction tuning LLMs.** **A2:** It appears there may have been some confusion regarding the datasets used in our work. To clarify: 1. **Task Mismatch:** The datasets MELD, IEMOCAP, and EmoryNLP are utilized by DialogueLLM for emotion recognition in conversations (ERC), focusing on the text modality. However, our work, Emotion-LLaMA, is fundamentally different as it emphasizes multimodal emotion recognition and reasoning by integrating audio, visual, and textual features to provide a far more comprehensive analysis of emotional states. 2. **Requirement Mismatch:** The instruction data used in DialogueLLM includes only emotion classification labels, lacking the depth and richness required for a nuanced understanding of emotions. Our MERR dataset, in contrast, provides instructions that include detailed multimodal emotional descriptions, enabling Emotion-LLaMA to analyze both external emotional expressions and internal psychological states with significantly greater accuracy. 3. **Modality Mismatch:** DialogueLLM’s input is limited to textual data, with no incorporation of audio or video modalities, as referenced in the DialogueLLM paper. This is a critical limitation that our approach directly addresses by incorporating comprehensive multimodal data into the instruction-tuning process, filling an important gap in current methodologies. --- **Q3: The selection procedure is still unclear - i.e., how was the dataset sampled from MER2023?** **A3:** We have previously addressed the question regarding the videos used in the MERR dataset in response to Q3-A3, where we provided detailed information about the video sources. We will now address the new question regarding the dataset sampling procedure from MER2023 and MERR: 1. The samples from the MER2023 dataset, which includes movies and TV series, were selected by cropping based on timestamps from the subtitles. The data collectors used a variety of open-source tools, such as YuNet and face.evoLVe, to ensure that each visual frame contained only one person. However, the dataset also includes a significant number of unlabeled samples that require further exploration. 2. The MERR dataset is derived from the MER2023 dataset. Initially, we filtered the samples based on the activation of specific Action Units to obtain coarse-grained descriptions. We then used LLaMA-3 to generate fine-grained descriptions, which were reviewed by emotional experts. This process resulted in 4,487 detailed annotated samples. --- **Q4 & Q5: No references are given for the supposed connection between AUs and emotion expression. It is unclear how stable this connection is in different contexts.** **A4 & A5:** The connection between Action Units (AUs) and emotion expression is not speculative; it is well-established and supported by extensive research: 1. **Validated by Research:** The Facial Action Coding System (FACS) has extensively validated the relationship between AUs and emotions, and this connection is considered stable and reliable across various contexts. This is a foundational aspect of emotion research. 2. **Practical Tools:** In practical applications, tools like OpenFace are widely used to extract AU information for precise emotion analysis, reinforcing the reliability of this connection in diverse scenarios. 3. **References:** To further substantiate this, we will include the following references in our revised manuscript: - [7] Facial Action Coding System, Environmental Psychology & Nonverbal Behavior, 1978. - [8] Affectiva-MIT Facial Expression Dataset (AM-FED): Naturalistic and Spontaneous Facial Expressions Collected In-the-Wild, CVPR 2013. - [9] Multi-Task Learning of Emotion Recognition and Facial Action Unit Detection with Adaptively Weights Sharing Network, ICIP 2019.

Authorsrebuttal2024-08-17

Detailed Response to Reviewer iNoC's Feedback (Part-2)

--- **Q6 & Q7: The authors say "Llama-3 refines annotations," but how this is exactly done and how the quality of this refinement can be assured is unclear. What is the protocol for human annotators? What are human "preferences" here?** **A6 & A7:** We appreciate your concern about the clarity of the LLaMA-3 refinement process. We would like to emphasize that the process is already thoroughly detailed in the manuscript, with specific evidence provided in several sections, including Table 9. 1. **Input Preparation:** The input preparation process is explicitly described, where we begin with coarse-grained multimodal descriptions derived from sources such as MiniGPT-v2, Action Units, Qwen-Audio, and Lexical Subtitles. This foundational step ensures that all relevant multimodal information is integrated before refinement. 2. **Prompt Construction:** The construction of prompts for LLaMA-3, as presented in Table 9, is designed to guide the model in generating detailed and coherent emotional descriptions. The prompt template shown in Table 9 is specifically adapted for the refinement task, ensuring that the process is aligned with the objectives of generating fine-grained annotations. 3. **LLaMA-3 Processing and Post-Processing:** The manuscript clearly explains how LLaMA-3 processes these prompts to produce refined descriptions that integrate information from all modalities, providing a nuanced interpretation of the emotional content. The subsequent post-processing step ensures consistency and removes any irrelevant information, further refining the outputs. 4. **Construction of Instruction-Following Data:** - **Initial Dataset and Automated Refinement:** The entire process, starting from the coarse-grained dataset of 28,618 samples and moving through automated refinement by LLaMA-3, is meticulously documented. This step is crucial for transforming basic annotations into detailed, instruction-following data. - **Normalization, Filtering, and Manual Review:** We have detailed how normalization and filtering techniques are applied to standardize the format of the refined descriptions and exclude inconsistent samples. Following this, four domain experts manually review the outputs, ensuring that only high-quality samples are included. The evaluation focuses on five critical points: visual modality accuracy, audio modality accuracy, textual modality accuracy, reasoning process correctness, and reasoning result correctness. 5. **Final Dataset:** The final dataset, consisting of 4,487 high-quality, fine-grained samples, is the result of a rigorous combination of automated refinement and expert human review. This process is clearly illustrated in the manuscript, ensuring transparency and reproducibility. **Supporting Evidence in the Manuscript:** - **Table 9**: The instruction template used by LLaMA-3 is provided, clearly demonstrating how the model is guided in generating refined annotations. - **Algorithm 1**: Outlines the overall annotation procedure, providing a comprehensive view of the steps involved. - **Supplementary Materials**: Additional details, including the flowchart of the refinement process, are included to further support the clarity of our approach. --- **Q9: Training details that are needed to understand the approach should be part of the main paper.** **A9:** We appreciate your focus on training details. The main paper includes essential information to understand our approach, but due to the complexity, not every detail could be included without affecting readability. For those needing more in-depth information, the full codebase is available in the anonymous repository, which covers all specifics, including hyperparameters and model configurations, necessary for replication. This ensures the paper remains clear and accessible, while the repository provides comprehensive details for those seeking further technical insights. ---

Authorsrebuttal2024-08-17

Detailed Response to Reviewer iNoC's Feedback (Part-3)

--- **+Q10: There should be a proper human evaluation of this process, at least on a part of the dataset.** **A10:** We conducted a thorough human evaluation of our annotation process to ensure the quality and relevance of the results. Specifically, we randomly selected 20 video samples from each of the nine emotion categories, resulting in 180 fine-grained annotations. These annotations were then randomly shuffled and evaluated by five volunteers. The evaluation criteria were carefully designed to assess: - Accuracy of the visual modality description - Accuracy of the audio modality description - Accuracy of the textual modality description - Correctness of the reasoning process - Correctness of the reasoning result Each volunteer scored the descriptions on a scale from 0 to 5, and the aggregated results are as follows: | Volunteer | Angry | Happy | Surprise | Fear | Sad | Worry | Neutral | Doubt | Contempt | Mean | |-----------|-------|-------|----------|------|-----|-------|---------|-------|----------|------| | V1 | 4.23 | 4.33 | 4.31 | 4.19 | 4.27| 4.33 | 4.04 | 4.55 | 4.79 | 4.34 | | V2 | 4.54 | 4.78 | 4.88 | 4.81 | 4.60| 4.62 | 4.65 | 4.55 | 4.68 | 4.67 | | V3 | 3.92 | 4.11 | 4.25 | 4.50 | 4.00| 4.24 | 3.78 | 4.27 | 4.05 | 4.12 | | V4 | 3.69 | 3.94 | 3.62 | 3.69 | 3.93| 4.19 | 3.39 | 4.00 | 4.53 | 3.90 | | V5 | 3.92 | 4.72 | 4.25 | 4.06 | 4.00| 4.19 | 4.17 | 4.32 | 4.58 | 4.26 | The overall average score of 4.258 demonstrates the high quality and logical alignment of our annotations. It’s noteworthy that the "Neutral" category received a slightly lower score, suggesting that when a character’s emotion is neutral, the cues are subtler, making automatic annotation more challenging. We acknowledge this limitation and plan to address it in future work. All details of the human evaluation, including the code and assessment results, are available in the anonymous repository for your review. --- **+Q12: The issue about internal emotional states that I raised is not mentioned in the rebuttal.** **A12:** We have indeed addressed the topics of displayed emotions and internal emotional states in our responses to Q6, Q7, and Q8. To reiterate, Emotion-LLaMA does not differentiate between displayed emotions and internal states; rather, it considers both aspects equally in its analysis. As demonstrated in Table 14 and Table 15, Emotion-LLaMA effectively identifies displayed emotions, such as a strong emotional tone or a furrowed brow, and accurately infers emotions like surprise and anger. Moreover, as shown in Table 4, even when a person in the video displays a prominent smile, Emotion-LLaMA can infer internal emotional states, such as annoyance, based on the spoken content, leading to the correct inference of anger. --- We kindly ask you to review our updated materials with these clarifications in mind. We are confident that our work provides a substantial contribution to the field, and we appreciate your thoughtful consideration of our paper. Sincerely, Authors

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC