Summary
This paper presents Orca to incorporate personality traits. Orca consists of four stages: inferring personality traits, augmenting data, constructing datasets, and modeling and training. The paper also introduces OrcaBench, a benchmark for evaluating LLM-generated content on social platforms. Experiments show that Orca achieves better performance on OrcaBench.
Strengths
1.The paper proposes a framework, Orca, that integrates personality traits into the data processing and training of custom LLM characters, addressing a gap in previous research.
2. The introduction of OrcaBench represents a significant contribution, providing a tool for evaluating the quality of LLM-generated content across multiple scales.
3. The experiments demonstrate Orca's superior performance on the OrcaBench benchmark.
Weaknesses
1.In terms of OrcaBench evaluation benchmark, the model is only compared to general open-source models. Experiments show that the proposed model achieves superior performance on this self-build benchmark. It is not enough to demonstrate its excellence of role-playing abilities. Maybe a human evaluation to assess the role-playing ability is more neutral.
2.BigFIVE (OCEAN) has its own questionnaire as a scale. Have you ever tried if your model role plays a certain personality and psychologists (or people with psychology background) conduct interviews, can it reflect a specific personality?
3.In this manuscript, there are several typos present, which significantly hinder the ease of reading. These mistakes can be particularly noticeable in certain paragraphs, such as the those included within the QUESTIONS. It appears that further refinement and careful proofreading are essential to enhance the overall readability and professionalism of the manuscript.
4.In Table 4, why does PTIT underperform PCIP with regards to CPR PTR PKR PSS by a large margin?
Questions
1.In line 287-288, “ The figure 3.3 illustrate the workflow of assistant follow the character and PCIP to generate psychological activities and response content”, I have not found figure 3.3.
2.in line 449-450, “enriching the information prior to the final output of content in a similar way to COT, as shown by the bold blue text in the bubble in Figure 3.3.”, I have not found Figure 3.3.
3.In line 445-446, “In contrast to the findings of Result 3 - PCIP-WPM, the addition of psychology activities did not result in a significant decrease in PSS scores compared between PTIT and”, what you do mean by “Result 3”. Maybe it’s the results in Table 3?
4.In Table 4, why does PTIT underperform PCIP regarding CPR PTR PKR PSS by a large margin?
5.In terms of OrcaBench evaluation benchmark, the model is only compared to general open-source models. Experiments show that the proposed model achieves superior performance on this self-build benchmark. It is not enough to demonstrate its excellence of role-playing abilities. Is there a human evaluation to assess the role-playing ability?
6.BigFIVE (OCEAN) has its own questionnaire as a scale. Have you ever tried if your model role plays a certain personality and psychologists (or people with a psychology background) conduct interviews, can it reflect a specific personality?