Large Language Models are Capable of Offering Cognitive Reappraisal, if Guided

Large language models (LLMs) have offered new opportunities for emotional support, and recent work has shown that they can produce empathic responses to people in distress. However, long-term mental well-being requires emotional self-regulation, where a one-time empathic response falls short. This work takes a first step by engaging with cognitive reappraisals, a strategy from psychology practitioners that uses language to targetedly change negative appraisals that an individual makes of the situation; such appraisals is known to sit at the root of human emotional experience. We hypothesize that psychologically grounded principles could enable such advanced psychology capabilities in LLMs, and design RESORT which consists of a series of reappraisal constitutions across multiple dimensions that can be used as LLM instructions. We conduct a first-of-its-kind expert evaluation (by clinical psychologists with M.S. or Ph.D. degrees) of an LLM's zero-shot ability to generate cognitive reappraisal responses to medium-length social media messages asking for support. This fine-grained evaluation showed that even LLMs at the 7B scale guided by RESORT are capable of generating empathic responses that can help users reappraise their situations.

Paper

References (79)

Scroll for more · 38 remaining

Similar papers

Reviewer ArmX6/10 · confidence 3/52024-05-09

Summary

Long-term mental well-being requires emotional self-regulation, starting by engaging with cognitive reappraisals that uses language to change negative appraisals that an individual makes of the situation. This paper hypothesizes that this task can be elicited from LLMs if they are guided by carefully crafted principles. Based on this, this paper introduces RESORT (REappraisals for emotional SuppORT), a psychologically-grounded framework that defines a constitution for a series of dimensions, motivated by the cognitive appraisal theories of emotions. The authors conduct an extensive evaluation of LLMs (GPT4-v, Llama2 13 B and Mistral 7B) for their cognitive reappraisal capability by clinical psychologists with M.S. or Ph.D. degrees, who judged LLM outputs (as well as human responses) in terms of their alignment to psychological principles, perceived empathy, as well as any harmfulness or factuality issues. Experimental results show that LLMs (even those at the 7B scale) produce cognitive reappraisals that significantly outperform human-written responses as well as non-appraisal-based prompting.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1. This paper is well-written and easy to follow. All the implementation details are clearly shown in the main paper, as well as the appendix. 2. The authors evaluate the performance by clinical psychologists with M.S. or PhD. Degrees, making the findings more sound. 3. The proposed method can significantly outperform baseline methods and even oracle responses.

Reasons to reject

1. The method is relatively too simple without too much technical contribution. 2. The scope and the focus of this paper, Cognitive Reappraisal, is too narrow and specific, limited the potential impact of this paper.

Questions to authors

1. I am interested in how can LLM response better than human oracle, written by PhD student in psychology. It will be better if the authors can show some case study, as well as the human annotation to provide some insight. 2. Also, providing a case study about the responses with and without appr and cons may help the reader to understand the effectiveness of each component. I find a case study in Table 6 in the Appendix. It will be better to mention it in the main paper. 3. The evaluation is too subjective and rely on expert knowledge, making it challenging to judge the effectiveness, even for the reviewer. For example, for the example in Figure 2, it is not very clear to me that how the guided response is better than unguided. 4. The examples in Table 9 are also a little bit confusing. It is not clear to me why the first response is Lack of Specific Guidelines / Actionable Steps.

Reviewer Zhu46/10 · confidence 3/52024-05-09

Summary

This paper explores the potential of Large Language Models (LLMs) to offer cognitive reappraisal, a psychological strategy that helps individuals change their negative appraisals of situations. The authors introduce the RESORT framework, which consists of reappraisal constitutions across multiple dimensions that can guide LLMs in generating empathic responses to support individuals in reappraising their situations. The study includes an expert evaluation by clinical psychologists, showing that LLMs guided by RESORT can produce effective cognitive reappraisal responses to social media messages seeking support.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1. The paper introduces a novel application of LLMs in providing cognitive reappraisal, showcasing the potential for advanced psychological capabilities of LLM. 2. They conducted an extensive evaluation of LLMs for their cognitive reappraisal capability.

Reasons to reject

1. In the experiment, Oracle responses were written by one person and may have subjective biases. In the experiment, it can be seen that the performance often deteriorates when +app+cons are used simultaneously. 2. LLM is sometimes sensitive to prompt words. Have you tested the performance of different Constructions prompt words. 3. Is there any other data source, such as real users asking questions when consulting with psychologists. Experiments from different data sources can demonstrate the generalization of the method.

Reviewer LP977/10 · confidence 4/52024-05-10

Summary

The paper introduces the RESORT framework to guide the large language models (LLMs) to generate reappraisals for emotional support, focusing on the six appraisal dimensions of emotions covering the Reddit posts. Both expert and automatic evaluations are conducted to understand the cognitive appraisals (reappraisals).

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

Pros: the paper may contribute a benchmark dataset to evaluate the large language models (LLMs) in generating (or reframing) reappraisals for emotional support. This contribution subjects to the release of the dataset to the research community.

Reasons to reject

Cons: C1: the evaluation samples are limited to 400 posts, 100 for each of the four domains. The oracle responses are 20 plus one highest up-voted comments. This might be relatively small to achieve a reliable evaluation. C2: the RESORT framework adopts a zero-shot setup to elicit responses. This may be due to the limited number of samples as indicated in the above C1, but in-context learning only requires several examples to prompt a (maybe better) response. Any explorations for these setups?

Questions to authors

Other comments and suggestions: How to determine the order of the six appraisal dimensions in Iterative Guided Refinement? Carol is a more common name than Erin. Andy, Betty, and Carol also sound harmony.

Reviewer WuPB5/10 · confidence 4/52024-05-24

Summary

The paper presents RESORT, a framework designed to guide large language models (LLMs) in generating cognitive reappraisal responses, a strategy in psychology aimed at changing negative appraisals to improve emotional well-being. The evaluation, conducted by clinical psychologists, demonstrates that even smaller-scale LLMs, when guided by RESORT, can produce effective cognitive reappraisals that significantly outperform human-written responses and non-appraisal-based prompts. The study highlights the potential of LLMs to assist in emotional support through targeted reappraisal, offering a scalable and efficient alternative to traditional methods.

Rating

5

Confidence

4

Ethics flag

1

Reasons to accept

1. Innovative Framework: The paper introduces a novel framework (RESORT) that integrates psychological principles to guide LLMs in generating cognitive reappraisals, addressing a crucial aspect of emotional self-regulation. 2. Expert Evaluation: The evaluation by clinical psychologists provides robust validation of the framework's effectiveness, enhancing the credibility of the findings.

Reasons to reject

1. Uncomprehensive Evaluation: The evaluation should include comparisons with more large language models such as GPT-40, Gemini, and domain-specific models like MeChat to provide a broader context of the agent’s performance. 2. Closed-Sourced Data: The evaluation is based on closed-sourced data, raising concerns about the reproducibility and transparency of the results. It would be beneficial to conduct evaluations on at least one open-sourced dataset. 3. Annotator Background Clarity: The paper lacks detailed information about the background of the human annotators. Moreover, Table 2 indicates low inter-annotator agreement (IAA) on certain metrics (ALGN and EMPT), which needs to be addressed. 4. Effectiveness of Iterative Guided Refinement: Results in Table 3 suggest that the Iterative Guided Refinement method does not significantly improve performance, questioning its utility in the proposed framework.

Questions to authors

Do you have ideas on how to improve the inter-annotator agreement (IAA) for human evaluations, particularly for the ALGN (alignment with reappraisal constitutions) and EMPT (empathy) metrics?

Ethics concerns details

N/A

Reviewer LP972024-06-05

Since the authors have put great efforts into this dataset, I hope it can be released to the research community for follow-up research papers.

Reviewer Zhu42024-06-05

I also hope the authors can release the dataset to the research community.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC