Summary
This paper dives into an interesting territory: the personalities of Large Language Models (LLMs). The researchers inquire whether LLMs possess discernible personalities and how these can be evaluated or even influenced. They introduce the Machine Personality Inventory (MPI), based on well-established personality trait theories and psychometric inventories, to evaluate and quantify the "personality" of LLMs. The team provides empirical evidence demonstrating the existence of personality-like tendencies in LLMs. They further propose a method, P2 (PERSONALITY PROMPTING), which can induce specific personalities in LLMs, thereby promoting diversity in their behaviors.
---
Rebuttal Acknowledgement: I appreciate the authors' response. They have addressed some of my concerns. I am thus raising my score from 4 to 5.
Strengths
(1) The idea of studying the personality of LLM is innovative and relatively unexplored in AI. Using the Big Five personality traits as a framework to evaluate and induce personality into LLMs is a novel application of psychological theory to machine learning.
(2)The creation of the Machine Personality Inventory (MPI), based on standard psychometric inventories, offers a systematic evaluation method for assessing the personalities of LLMs. This helps fill a gap in current AI research, which has focused mainly on abstract visual reasoning.
(3) The use of a multiple-choice question-answering suite to quantitatively evaluate LLMs' behaviors from a personality perspective is a major strength of this paper. The use of mean and standard deviation measurements provides a clear and rigorous means of assessing the LLMs' personality traits.
(4) The successful implementation of the PERSONALITY PROMPTING (P2) method to control the LLMs' behavior shows the practical application of the research. The work has potential real-world applications in creating more relatable and human-like AI systems. The paper explores the existence of personality in LLMs and validates the possibility of inducing different personalities into LLMs.
Weaknesses
(1) Although the Big Five personality traits provide a valuable framework for assessing personality, human personality is complex and multifaceted, and this model may not fully capture the nuances of personality.
(2) The paper seems to assume that LLMs inherently possess a certain personality, which might be a simplification. It could be argued that any perceived personality in an LLM is a projection of the human users or an artifact of the data on which the LLM was trained.
(3) The selection of LLMs was based on some prerequisites, which might have limited the diversity of the models evaluated. The study primarily used GPT-3.5 for inducing personality due to its similarity to human statistics, which might limit the generalizability of the findings.
(4) While the work has potential real-world applications, there's also a risk of misuse. For instance, AI systems could potentially manipulate people by mimicking desirable personality traits. An ethical discussion around this is needed.
(5) The process of inducing personality involves the generation of a 'naive' natural language prompt, which could be subject to human bias. The system's 'personality' might then reflect the biases of the person who created the prompt rather than being an independent characteristic of the LLM itself.
(6) The paper uses human evaluators from Prolific to determine if the generated responses correspond to the induced personality. The reliability of this method might be questioned as it depends on non-expert human judgment (i.e., nonpsychologists), which can be very unreliable as this is not as easy task.
Questions
(1) The authors propose the Machine Personality Inventory (MPI) as an evaluation tool for machine personality. However, it isn't entirely clear how the MPI translates the complexity of human personality into a machine-evaluable format. Could the authors expand on this aspect?
(2) The authors utilized GPT-3.5 for the study because of its similarity to human statistics and superior performance in various natural language tasks. However, have the authors considered the performance of other models and how the P2 method might affect them?
(3) The authors presented that they could control the five personality factors in LLMs. Can the authors also control the intensity of these traits, or are the traits presented as binary options (present or absent)?
(4) Can the authors elaborate on the practical use cases of LLMs induced with specific personalities? How might this enhance or deter interactions between LLMs and human users in real-world applications?
(5) The paper indicated that LLMs behaved like persons with personalities, matching corresponding human-like behaviors. Could the authors elaborate on the criteria used to draw this comparison? What aspects of human behavior were considered when concluding that LLMs possess personality?
(6) See weakness #4 about the risk of misuse and comment, please.
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
While the authors acknowledge some of the potential downsides of their work, they could have elaborated more on these points:
The authors discuss the introduction of personality to LLMs, but they could explore more potential negative societal impacts. For instance, could this personality induction lead to manipulation or the creation of overly persuasive AI systems? Could introducing certain personalities contribute to existing biases in AI systems, and how might this be mitigated?
The authors based their research on the Big Five personality factors, which is a broadly accepted model of human personality, but it's not without its limitations and critics. An exploration of these potential limitations and how they could influence the results would provide a more balanced perspective.
There's a need to explore the limitations of the P2 method in inducing personality. While it's a promising approach, there could be limitations such as lack of subtlety in personality induction or inability to induce multiple traits simultaneously.
It would also be beneficial to discuss the limitations of the models used in the study. Each AI model may have specific limitations that could affect the results of the personality induction.