Idiographic Personality Gaussian Process for Psychological Assessment

We develop a novel measurement framework based on a Gaussian process coregionalization model to address a long-lasting debate in psychometrics: whether psychological features like personality share a common structure across the population, vary uniquely for individuals, or some combination. We propose the idiographic personality Gaussian process (IPGP) framework, an intermediate model that accommodates both shared trait structure across a population and"idiographic"deviations for individuals. IPGP leverages the Gaussian process coregionalization model to handle the grouped nature of battery responses, but adjusted to non-Gaussian ordinal data. We further exploit stochastic variational inference for efficient latent factor estimation required for idiographic modeling at scale. Using synthetic and real data, we show that IPGP improves both prediction of actual responses and estimation of individualized factor structures relative to existing benchmarks. In a third study, we show that IPGP also identifies unique clusters of personality taxonomies in real-world data, displaying great potential in advancing individualized approaches to psychological diagnosis and treatment.

Paper

References (58)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer A4W77/10 · confidence 4/52024-07-08

Summary

This paper presents a multi-output Gaussian process model for classification in the context of psychological studies. It's based on a linear model of co-regionalisation leveraging unit-level latent factors and RBF kernels to jointly model population and individual personality traits to tackle an old debate in psychometrics regarding their original nature. The authors leverage a variational formulation to derive a lower bound for the model evidence and optimise the individual-specific, population, and GP-related parameters. An extensive simulation and real-data study is proposed, presenting longitudinal and multi-subject survey analyses, personality correlations and predictions.

Strengths

This paper is remarkably well-written and presented. The flow of derivations is easy to follow, and the illustrations are helpful and of excellent quality. The treated problem is far from trivial, with data that present several degrees of correlations (time, individual, population) and technical constraints (missing and categorical data). The experiments are impressive, with extensive comparison against many competitors and excellent results overall. The method shows great promise in providing practical insight into the domain, and even though I'm not a specialist in psychology, the ability to reconcile idiographic and nomothetic approaches seems particularly valuable.

Weaknesses

One could argue that the methodological novelty is limited as the presented multi-output GP model is fairly well-known in the GP community. However, I think its similarities and differences with GPLVM and GPDM are well presented, and the overall application is far from trivial regarding the nature of data (categorical outputs, latent variables, missing data, ...). This relative weakness is more than compensated by the strength of results and the potential such a methodology brings in a field like psychology, where the signal-on-noise ratio is generally low, and the nature of measurements is often challenging.

Questions

I only have a few questions regarding computation times, scaling and robustness of the inter-output covariance matrices, which are well-known limitations of co-regionalisation GP models in practice. - In this sense, could you provide numerical evidence and comparison regarding training and prediction times for IPGP and competitors? - Even though the number of tasks/topics remains probably limited in psychological applications, could you discuss a bit more, in Section 4.1 **Result**, the decrease in performance between full IPGP and the low-rank version. What could we expect for higher dimensions? Would the low-rank approximation remain robust for a more significant gap with the actual rank? - In my experience, several multi-output GP models (or maybe their implementations) are somewhat unstable when it comes to estimating $\textbf{K}_{task}$, and pathological cases can often arise. Have you experienced such problems? And if so, how did you manage to tackle this issue? **Typos:** Page 5: The indices of the CMD definition have a slight error. All terms are written as $R_1$, and the norm is 'f' instead of '$l_2$' Page 6: The title of 4.2 "Cross-secitonal" ==> Cross-sectional Page 7: "Both correlation matrices **displace** a block pattern" ==> display? Page 7: "questions corresponding negative emotionality" ==> corresponding to

Rating

7

Confidence

4

Soundness

4

Presentation

4

Contribution

3

Limitations

In my opinion, the limitations are adequately discussed. The application spectrum and position in the related literature are well documented.

Reviewer A4W72024-08-13

Thank you for your thorough answers. I understand the technical constraints coming from such costly and time-consuming studies involving human subjects. While I acknowledge that this discussion about running time might be superfluous for your application, I always prefer to see beforehand 'how much exactly' it will cost me when I intend to test a method on my data, and I suspect that future readers would, too. This is especially true for methods like MTGP, which I know are coming at a not negligible cost. That being said, I commend the authors' efforts to provide additional evidence and thus improve the confidence one can grant this paper. More generally, I support the authors' view about 'building bridges' between ML methodologies and their applications. Although I consider myself more of a methodological researcher, I notice that we too often focus on 'novelty' by principle while neglecting meticulous and well-conducted applications. Such studies are crucial for advancing more experimental disciplines like psychology with all the rigour it deserves from ML researchers. We all agree that this paper aims not to *advance ML* vastly, but I appreciate that it *uses ML to advance*.

Authorsrebuttal2024-08-13

> I always prefer to see beforehand 'how much exactly' it will cost me when I intend to test a method on my data, and I suspect that future readers would, too. This is especially true for methods like MTGP, which I know are coming at a not negligible cost. This view resonates with us. We're more than happy to add this discussion and these results regarding running time to the paper to help complete the story. We appreciate your original comment and the invitation to run these additional experiments. > That being said, I commend the authors' efforts to provide additional evidence and thus improve the confidence one can grant this paper. More generally, I support the authors' view about 'building bridges' between ML methodologies and their applications. Although I consider myself more of a methodological researcher, I notice that we too often focus on 'novelty' by principle while neglecting meticulous and well-conducted applications. Such studies are crucial for advancing more experimental disciplines like psychology with all the rigour it deserves from ML researchers. We all agree that this paper aims not to advance ML vastly, but I appreciate that it uses ML to advance. We truly appreciate your support!

Reviewer mvMV6/10 · confidence 5/52024-07-10

Summary

Gives a multi-task/output GP formulation for multiple time-series (or intrinsic co-regionalisation model) for pyschological assessments. Design of the factor loadings informs the task-correlations and reflects the knowledge of individual's correlation between his responses and inter-person correlations. More of an application paper rather than a paper advancing machine learning; but a good application nonetheless.

Strengths

1. The application is a natural fit for the model chosen, and results are good and convincing. 2. The use of Bayes factor for model testing make the paper better as an reference for this application. 3. Sections 4.2 and 4.3 evaluates the model from different aspects.

Weaknesses

1. The paper is weak from the perspective of advancing state-of-the-art machine learning algorithms. 2. The paper is totally unrelated to GPLVM and tasks that GPLVM are designed for. Mentions of GPLVM only confuses the reader. 3. Equation 3 needs fixing. 4. In section 3.2, the method is variational inference, and not stochastic variational inference. 5. Description of the setup in section 4.1 needs to be clearer to explain also in terms of $K_{task}$. 6. Line 192 mentions "informative prior". What is this "informative prior"? *Minor* 7. Line 319: "addressing" is too strong. Suggest to change to "contributing". 8. Line 321: "than" -> "over"

Questions

I do not have any critical questions.

Rating

6

Confidence

5

Soundness

3

Presentation

3

Contribution

2

Limitations

Yes.

Reviewer YTHZ3/10 · confidence 5/52024-07-12

Summary

1. This paper introduces an innovative measurement framework utilizing the Gaussian process coregionalization model to resolve the question of whether psychological attributes such as personality exhibit a universal structure among the populace or are uniquely individualized. 2. An Idiographic Personality Gaussian Process (IPGP), a hybrid model that accounts for both the commonality of traits across people and individual-specific "idiographic" variations. 3. Gaussian process coregionalization model to interpret the responses from grouped survey batteries, adapted for non-Gaussian ordinal data, and employs stochastic variational inference for estimating latent factors. 4. The application of IPGP on both synthetic data and an original survey demonstrates its performance.

Strengths

1. The paper’s main innovation lies in its unique combination of multitask Gaussian process, which results in a conceptualization of the subject matter. 2. The innovative use of multitask Gaussian process in this paper offers a promising solution to a long-standing challenge of psychological assessment. 3. The quantitative and qualitative methods in this paper provide a potential in advancing psychological diagnosis and treatment. 4. The interdisciplinary nature of the research is interesting. 5. The authors provide a holistic view of the issue at hand.

Weaknesses

1. there are many MTGPs, the author ignored comparison with them. 2. incorrect description, such as: "multi-task structure is also known as the linear model of coregionalization (LMC)". 3. This represents a particular implementation of MTGP, but it doesn't introduce any novel methodological advancements. 4. The paper presents a valuable contribution, but its innovative aspects are limited, as it largely builds upon existing theories without introducing significant new insights. 5. The literature review appears somewhat limited. Expanding it to include more recent or relevant MTGPs could provide a more comprehensive context for the research. 6. The statistical analysis used in the study seems inadequate given the complexity of the data. 7. The research seems to be more of an incremental advancement rather than a novel contribution, which may limit its impact on the field. 8. The experimental evidence provided in the paper is not as comprehensive as it should be to support the claims made, suggesting a need for more extensive testing. 9. Only one experiment is related to psychological assessment. 10. No idiographic personality is discovered and presented in the paper.

Questions

Please see the weakness above

Rating

3

Confidence

5

Soundness

2

Presentation

3

Contribution

2

Limitations

Please see the weakness above

Reviewer YTHZ2024-08-08

Employing other advanced MTGPs for this psychological task is straightforward. However, the author intentionally abandoned more comparisons with other MTGPs. The innovation of the method proposed in this paper is quite limited. This article makes a very small contribution to the field of Gaussian processes Compared to state-of-the-art methods, its advancements and advantages are minimal. As this is an applied work, I suggest the authors submit their paper to a journal focused on bioinformatics.

Authorsrebuttal2024-08-12

Thank you for your response, although we respectfully disagree. We strongly believe work such as ours has a place at this conference. > As this is an applied work, I suggest the authors submit their paper to a journal focused on bioinformatics. The call for papers explicitly encourages submissions of applications and highlights the participation of diverse communities beyond core ML: > [NeurIPS] brings together researchers in machine learning, neuroscience, statistics, optimization, computer vision, natural language processing, life sciences, natural sciences, social sciences, and other adjacent fields. > We invite submissions presenting new and original research on topics including but not limited to the following: > - Applications > [...] > - Machine learning for sciences (e.g. climate, health, life sciences, physics, social sciences) [...] We do not believe it is in the spirit of the call for papers to dismiss applied work out of hand. > The innovation of the method proposed in this paper is quite limited. > This article makes a very small contribution to the field of Gaussian processes The reviewer guidelines explicitly encourage diversity in contributions beyond purely methodological improvements to (MT)GPs: > There are many examples of contributions that warrant publication at NeurIPS. These contributions may be theoretical, methodological, algorithmic, empirical, connecting ideas in disparate fields (“bridge papers”), or providing a critical analysis. Our contributions here required both considerable expertise in GP modeling (including the expertise required to build several necessary innovations highlighted by reviewer A4W7) and considerable and sustained engagement with another domain to ensure success. Our strong results reflect a sophisticated model shedding light on an important psychological question. We do not believe it is in the spirit of these guidelines to dismiss our non-methodological (empirical, bridging communities mentioned in the CFP, etc.) contributions out of hand -- especially as they are supported by methodological contributions as well!

Reviewer mvMV2024-08-10

I think it is a matter of interpreting "bridge". It is clearly a "bridge" in terms of bringing new applications in. It is not a "bridge" in terms of bringing new ideas to advance ML --- an example of which is Random Matrix Theory. For GPLVM, I am agreeable to it being mentioned in related work. Claiming in the introduction that you "advances on the ... GPLVM" is simply too much for me to take. I may reconsider the score during the reviewer discussion stage.

Authorsrebuttal2024-08-10

Thank you for your response! > Claiming in the introduction that you "advances on the ... GPLVM" is simply too much for me to take. We are happy to rephrase this passage, in particular to avoid the acronym GPLVM entirely. We intended to indicate that our model was among a family of spiritually related models incorporating latent variables and Gaussian processes in their construction (it is "a" latent variable GP model), not that we were advancing "the" famous GPLVM model from Lawrence.

Authorsrebuttal2024-08-13

Dear Reviewer f9td, Thank you for your feedback. We understand your concerns about selection bias and its impact on our results, though we believe the model still holds value. Here is our response: Social and personality psychology often faces challenges with model building due to reliance on non-representative samples, typically from student populations. We follow current standards by aiming to broaden our sample within budget and available diversity. We will clearly state these limitations in the manuscript and suggest how future studies can test the model's applicability across different populations and cultures. We will also include a more direct discussion on the diversity of our sample. For the first case study, the dataset from Soto (2019) was relatively diverse, using Qualtrics and quota sampling to ensure representative samples of the U.S. population: * age: 11% ages 18-24, 18% ages 25-34, 17% ages 35-44, 19% ages 45-54, 17% ages 55-64, 18% ages 65 and older * sex: 52% female, 48% male * race/ethnicity: 74% non-Hispanic white/Caucasian, 11% black/African American, 10% Hispanic/Latino, 3% Asian/Asian American, 2% American Indian/Native American * educational attainment: 10% did not complete high school, 33% high school graduate, 28% some college, 19% college graduate, 10% graduate or professional degree * annual household income (\$): 14\% <20,000, 12% 20,000-29,999, 11\% 30,000-39,999, 15\% 40,000-49,999, 26\% 50,000-79,999, 22\% 80,000+ In the second study, participants in our longitudinal study (the third experiment) are mostly college students with an average age of 20.23 (SD = 1.94); 70% are female, 26% are male and 4% are self-identified as other; ethnicities self-reported as 42% Caucasian, 39% Asian, 12% African American, and 7% Other. We used the BFI-2 [Soto & John, 2007], a widely used personality measure across cultures and ages [Yoshino et al., 2022; Beatrice et al., 2024] (see supplementary materials). We acknowledge the typical biases of convenient sampling in higher education, including socio-economic and ethnic diversity limitations, which we will explain more clearly in the manuscript.

Reviewer 5m3j6/10 · confidence 4/52024-08-15

Summary

UPDATE: I am updating my scores in light of the excellent authors' response. This paper considers psychometric data composed of ordinal responses $y_{ijt}$, each of which represent how unit _i_ answered survey item _j_ during time period _t_. The key characteristic of this item-response data is that the same $N$ units are longitudinally surveyed on the same $J$ items over $T$ periods. The paper takes an ordered logit factor modeling approach to analyzing such data, wherein $f_j^{(i)}(t) = \mathbf{w}_j^\top \mathbf{x}_i(t)$ is the latent ideal point of unit _i_ on on item _i_ at time _t_, and $\mathbf{w}_j$ and $ \mathbf{x}_i(t)$ are the $K$-dimensional loadings and factor vectors, respectively. The paper moreover advocates an "idiographic approach [which] emphasizes _intrapersonal_ variation by requiring distinct loadings $\mathbf{w}_j^{(i)}$" which are different for each unit $i$. It is able to effectively achieve this by exploiting the repeated measurements of each unit-item $(i,j)$ pair over different time periods $t$. The proposed model is an instance of a multi-task Gaussian process (MTGP). Intrapersonal variation is modeled using a unit-specific kernel $K_{time}^{i}$. The full covariance is then $JT \times JT$, resulting from a Kronecker product of $\mathbf{K}_{\textrm{time}}^{(i)}$ (applied to input $T$ periods) with a low-rank matrix $J \times J$ matrix that represents covariance between survey items (or "tasks"). This all extends the previously-presented linear model of coregionalization (LMC), which was originally developed for the simpler case where observations of $(i,j)$ pairs are not repeated over time. The paper derives a variational inference inference algorithm for the model, which follows closely from previous work. The paper then reports three sets of experiments: 1) a synthetic study that reports excellent results on parameter recovery (when the true parameters are known), 2) a re-analysis study that purports to show that a non-dynamic (i.e., non-idiographic) version of proposed model is able to identify the correct factor structure (K=5) from survey data designed to assess the "Big Five" psychometric traits, and 3) an illustrative case study involving a novel longitudinal data set.

Strengths

The paper is very well-written. It brings an interesting application to life and convinces the reader that the proposed modeling approach is well-tailored to the problem at hand. The paper gives a clear review of the prerequisite concepts and presents its modeling approach clearly. The paper covers related work in psychometrics well is convincing that the modeling approach is novel within that applied community. Experiments bear out this claim, as the proposed model performs much better than models which are currently used in psychometrics. The paper presents a novel longitudinal survey dataset composed of $N=93$ subjects who were given personality assessment surveys over the course of three weeks. If I am reading the paper correctly, each subject was asked to complete a survey _six times per day_. This data sounds highly non-trivial to collect, and should be considered a main contribution in itself.

Weaknesses

The paper is vague about its technical contributions. It makes the following hedged novelty statement: "the first multi-task GP latent variables model _for dynamic idiographic assessment_". I took this to mean that the proposed approach is not very new technically, but it has never yet been applied to dynamic idiographic assessments. However, the paper also makes the general claim that it "advances the literatures on Gaussian process latent variable models...". I am of the opinion that tailoring existing modeling frameworks to new applications does provide a technical contribution, as it contributes a new view and set of interpretations/metaphors that can help to better understand the abstract model. I think the paper does contribute in this way. But I am unsure if the paper further contributes more substantially to the area of Gaussian process latent variable models, as the paper is vague about that. There is sloppiness in some of the main equations. For instance, equation 2 confuses a distribution with a random variable stating $p(\mathbf{f}^{(i}) \sim \cdots$, and I believe equation 3 confuses an inner with an outer product (shouldn't it be $w_i w_i^\top$?). There's also some confusing overloading of symbols between the background and the model sections. The background sets up the idea that an ideographic approach has distinct loadings $w_j^{(i)}$ different across $i$. But the proposed model itself does not seem to sport this as each $w_j$ is global for each task. I believe effectively the model does achieve an idiographic interpretation through the unit-specific kernel, but the connection between the background and model section is not clear. I think the second experiment has some major flaws. The paper claims this experiment validates the "Big Five" because model performance peaks at $K=5$. However, the model is only fit for $K=1...5$. We can clearly see from the log likelihood numbers that model performance for IPGP is monotonically increasing in $K$ for the values considered, suggesting that it will likely be even better at larger values of $K$; if that were true, it would not validate the "Big Five" theory. As this is an applied paper, I would expect a much higher level of rigor on one of the two applied case studies.

Questions

As this is an emergency review, after the reviewer discussion period, I am not sure if the authors will be able to reply. But if they are able, I would just ask that they respond to the various points made in previous sub-sections.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

This paper did involve a fairly intense human subjects experiment for data collection, but the paper reports it was IRB-approved. Psychometrics of this kind are fairly common; I do not think the paper introduces any new potential negative societal impacts.

Program Chairs2024-08-16

Hi all, We re-opened the discussion for this paper to Aug 17 11:59pm ET for authors to reply to the emergency review. The review is pasted below. ====== **Summary** This paper considers psychometric data composed of ordinal responses $y_{ijt}$, each of which represent how unit _i_ answered survey item _j_ during time period _t_. The key characteristic of this item-response data is that the same $N$ units are longitudinally surveyed on the same $J$ items over $T$ periods. The paper takes an ordered logit factor modeling approach to analyzing such data, wherein $f_j^{(i)}(t) = \mathbf{w}_j^\top \mathbf{x}_i(t)$ is the latent ideal point of unit _i_ on on item _i_ at time _t_, and $\mathbf{w}_j$ and $ \mathbf{x}_i(t)$ are the $K$-dimensional loadings and factor vectors, respectively. The paper moreover advocates an "idiographic approach [which] emphasizes _intrapersonal_ variation by requiring distinct loadings $\mathbf{w}_j^{(i)}$" which are different for each unit $i$. It is able to effectively achieve this by exploiting the repeated measurements of each unit-item $(i,j)$ pair over different time periods $t$. The proposed model is an instance of a multi-task Gaussian process (MTGP). Intrapersonal variation is modeled using a unit-specific kernel $K_{time}^{i}$. The full covariance is then $JT \times JT$, resulting from a Kronecker product of $\mathbf{K}_{\textrm{time}}^{(i)}$ (applied to input $T$ periods) with a low-rank matrix $J \times J$ matrix that represents covariance between survey items (or "tasks"). This all extends the previously-presented linear model of coregionalization (LMC), which was originally developed for the simpler case where observations of $(i,j)$ pairs are not repeated over time. The paper derives a variational inference inference algorithm for the model, which follows closely from previous work. The paper then reports three sets of experiments: 1) a synthetic study that reports excellent results on parameter recovery (when the true parameters are known), 2) a re-analysis study that purports to show that a non-dynamic (i.e., non-idiographic) version of proposed model is able to identify the correct factor structure (K=5) from survey data designed to assess the "Big Five" psychometric traits, and 3) an illustrative case study involving a novel longitudinal data set. **Soundness**: 3: good **Presentation**: 3: good **Contribution**: 2: fair **Strengths** The paper is very well-written. It brings an interesting application to life and convinces the reader that the proposed modeling approach is well-tailored to the problem at hand. The paper gives a clear review of the prerequisite concepts and presents its modeling approach clearly. The paper covers related work in psychometrics well is convincing that the modeling approach is novel within that applied community. Experiments bear out this claim, as the proposed model performs much better than models which are currently used in psychometrics. The paper presents a novel longitudinal survey dataset composed of $N=93$ subjects who were given personality assessment surveys over the course of three weeks. If I am reading the paper correctly, each subject was asked to complete a survey _six times per day_. This data sounds highly non-trivial to collect, and should be considered a main contribution in itself. **Questions** As this is an emergency review, after the reviewer discussion period, I am not sure if the authors will be able to reply. But if they are able, I would just ask that they respond to the various points made in previous sub-sections. **Limitations** This paper did involve a fairly intense human subjects experiment for data collection, but the paper reports it was IRB-approved. Psychometrics of this kind are fairly common; I do not think the paper introduces any new potential negative societal impacts. **Flag For Ethics Review**: No ethics review needed. **Rating**: 4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly. **Confidence**: 4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Program Chairs2024-08-16

**Weaknesses** The paper is vague about its technical contributions. It makes the following hedged novelty statement: "the first multi-task GP latent variables model _for dynamic idiographic assessment_". I took this to mean that the proposed approach is not very new technically, but it has never yet been applied to dynamic idiographic assessments. However, the paper also makes the general claim that it "advances the literatures on Gaussian process latent variable models...". I am of the opinion that tailoring existing modeling frameworks to new applications does provide a technical contribution, as it contributes a new view and set of interpretations/metaphors that can help to better understand the abstract model. I think the paper does contribute in this way. But I am unsure if the paper further contributes more substantially to the area of Gaussian process latent variable models, as the paper is vague about that. There is sloppiness in some of the main equations. For instance, equation 2 confuses a distribution with a random variable stating $p(\mathbf{f}^{(i}) \sim \cdots$, and I believe equation 3 confuses an inner with an outer product (shouldn't it be $w_i w_i^\top$?). There's also some confusing overloading of symbols between the background and the model sections. The background sets up the idea that an ideographic approach has distinct loadings $w_j^{(i)}$ different across $i$. But the proposed model itself does not seem to sport this as each $w_j$ is global for each task. I believe effectively the model does achieve an idiographic interpretation through the unit-specific kernel, but the connection between the background and model section is not clear. I think the second experiment has some major flaws. The paper claims this experiment validates the "Big Five" because model performance peaks at $K=5$. However, the model is only fit for $K=1...5$. We can clearly see from the log likelihood numbers that model performance for IPGP is monotonically increasing in $K$ for the values considered, suggesting that it will likely be even better at larger values of $K$; if that were true, it would not validate the "Big Five" theory. As this is an applied paper, I would expect a much higher level of rigor on one of the two applied case studies.

Authorsrebuttal2024-08-16

Dear Reviewer 5m3j, Thanks for your valuable feedback. Please see our response below. > If I am reading the paper correctly, each subject was asked to complete a survey six times per day. Yes, that's correct. Collecting this data was a significant undertaking, and to our knowledge, we are among the very few groups that have ever collected data of such scope. > The paper is vague about its technical contributions. I am unsure if the paper further contributes more substantially to the area of Gaussian process latent variable models. We describe our model as a "latent variable GP model" since it belongs to a broader category of models involving latent variables and GPs (including but not exclusive to the famous GPLVM from Lawrence). Technically, our model also differs with GPLVM in that: (1) optimizes the factor loading matrix while marginalizing the latent variables, (2) accommodates categorical data through a non-Gaussian ordered logistic likelihood. Please see our other rebuttals for more clarification on our contributions. > Equation 2 confuses a distribution with a random variable stating $𝑝(𝑓^{(𝑖)})$, and I believe equation 3 confuses an inner with an outer product (shouldn’t it be $𝑤_𝑖𝑤_𝑖^⊤$?). Thank you for pointing out the typo in eq 2. We will fix it in revision. Regarding eq 3, please see our discussion of updated notations below. > The background sets up the idea that an ideographic approach has distinct loadings $𝑤_j^{(𝑖)}$ different across 𝑖. But the proposed model itself does not seem to sport this as each $𝑤_𝑗$ is global for each task. Thank you for bringing this up. We agree the notation could be improved -- the confusion here might come from overloading the use of variable w depending on its indexing. We will rename $w_i$ in Figure 1 and eq 3 to $Z^{(i)}$. The full updated notation is listed below: - Throughout our notation, superscript $(i)$ indicates unit and underscript $j$ indicates task. - $W_{\text{pop}}$ represents the K x J (K latent dimensions, J tasks) shared interpersonal loading matrix. - $Z^{(i)}$ represents the $K^*$ by J unit-specific low-rank loading matrix that serves to be the additional idiographic component, independent of $W_{\text{pop}}$. The rank $K^*$ of $Z^{(i)}$ need not be the same as the rank K of $W_{\text{pop}}$, and in fact we used $K^*=1$ in our experiments (see below). - $K_{\text{task}}^{(i)}$ is the unit-specific task covariance matrix with shared component $W_{\text{pop}}'$ $W_{\text{pop}}$ and unit-specific deviations ${Z^{(i)}}' Z^{(i)}$ of $K^*<K$ (with this revised notation the transpose in equation (3) is correct as written). Now $W_{\text{pop}}$ is a global parameter estimated for the entire population while $Z^{(i)}$s are unit-level parameters. These are combined in eq (3) to create a unique kernel for each unit reflecting both components. To aid identification and performance, we focused on $K^*=1$ in the experiments; we also tried $K^*=2$ only to find degraded performance but extra computational costs. Note that there is much less data to estimate unit-level deviations than there is to estimate population-level structure, so in general we might expect to take $K^* < K$ when modeling. However, the reviewer is correct that even in this rank-1 case, our setup still induces a unique $K_{\text{task}}^{(i)}$ as specified in eq 3. We will clarify the discussion here and include a list of notations in the appendix. > The paper claims this experiment validates the “Big Five” because model performance peaks at 𝐾=5. However, the model is only fit for 𝐾=1...5. Thank you for this valuable comment; we ran two additional factor analysis experiments accordingly. We first examine the performance of IPGP with higher model ranks $K \ge 5$ in the second case study (LOOPR); the results are shown in the table below. The results consistently support the model with rank precisely 5, which has both higher model evidence and lower BIC, providing much more convincing evidence for the Big Five theory than originally presented. In particular, the BIC has a decreasing trajectory from lower rank to 5 and an increasing one from 5 to higher rank, suggesting that increasing the rank to 5 is necessary for the model to capture the ideal structure of the data, but that exceeding this rank only overfits. **Performance of IPGP with ranks from 1 to 10 for LOOPR.** | RANK | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | | - | - | - | - | - | - | - | - | - | - | - | | **LL/N** ↑ | -1.478 | -1.477 | -1.477 | -1.477 | **-1.476** | -1.477 | -1.477 | -1.477 | -1.478 | -1.477 | | **BIC** (×10^11) ↓ | 1.2736 | 1.2726 | 1.2728 | 1.2726 | **1.2722** | 1.2725 | 1.2726 | 1.2732 | 1.2732 | 1.2734 | Please also see the related additional study described in our rebuttal to reviewer A4W7, where we varied the model rank among {2, 5, 8} in our simulation study (the true rank is 5). We saw similar results there, where the true rank consistently yielded the best model fit.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC