Summary
Long-term mental well-being requires emotional self-regulation, starting by engaging with cognitive reappraisals that uses language to change negative appraisals that an individual makes of the situation. This paper hypothesizes that this task can be elicited from LLMs if they are guided by carefully crafted principles. Based on this, this paper introduces RESORT (REappraisals for emotional SuppORT), a psychologically-grounded framework that defines a constitution for a series of dimensions, motivated by the cognitive appraisal theories of emotions.
The authors conduct an extensive evaluation of LLMs (GPT4-v, Llama2 13 B and Mistral 7B) for their cognitive reappraisal capability by clinical psychologists with M.S. or Ph.D. degrees, who judged LLM outputs (as well as human responses) in terms of their alignment to psychological principles, perceived empathy, as well as any harmfulness or factuality issues. Experimental results show that LLMs (even those at the 7B scale) produce cognitive reappraisals that significantly outperform human-written responses as well as non-appraisal-based prompting.
Reasons to accept
1. This paper is well-written and easy to follow. All the implementation details are clearly shown in the main paper, as well as the appendix.
2. The authors evaluate the performance by clinical psychologists with M.S. or PhD. Degrees, making the findings more sound.
3. The proposed method can significantly outperform baseline methods and even oracle responses.
Reasons to reject
1. The method is relatively too simple without too much technical contribution.
2. The scope and the focus of this paper, Cognitive Reappraisal, is too narrow and specific, limited the potential impact of this paper.
Questions to authors
1. I am interested in how can LLM response better than human oracle, written by PhD student in psychology. It will be better if the authors can show some case study, as well as the human annotation to provide some insight.
2. Also, providing a case study about the responses with and without appr and cons may help the reader to understand the effectiveness of each component. I find a case study in Table 6 in the Appendix. It will be better to mention it in the main paper.
3. The evaluation is too subjective and rely on expert knowledge, making it challenging to judge the effectiveness, even for the reviewer. For example, for the example in Figure 2, it is not very clear to me that how the guided response is better than unguided.
4. The examples in Table 9 are also a little bit confusing. It is not clear to me why the first response is Lack of Specific Guidelines / Actionable Steps.