Learning Social Welfare Functions

Is it possible to understand or imitate a policy maker's rationale by looking at past decisions they made? We formalize this question as the problem of learning social welfare functions belonging to the well-studied family of power mean functions. We focus on two learning tasks; in the first, the input is vectors of utilities of an action (decision or policy) for individuals in a group and their associated social welfare as judged by a policy maker, whereas in the second, the input is pairwise comparisons between the welfares associated with a given pair of utility vectors. We show that power mean functions are learnable with polynomial sample complexity in both cases, even if the comparisons are social welfare information is noisy. Finally, we design practical algorithms for these tasks and evaluate their performance.

Paper

Similar papers

Peer review

Reviewer 1QjM7/10 · confidence 4/52024-06-30

Summary

The authors study the learnability of social welfare functions given decisions data by a central decision-maker that is taking into account their constituents' welfare. They discuss PAC bounds according to a number of settings with a focus on weighted power mean functions. These settings include cardinal utility vectors under a target social welfare functions and pairwise comparisons between utility vectors. Learning is taken either without noise, with iid noise, or logistic noise. The authors validate their theoretical findings by learning welfare functions on proprietary data by Lee et. al. (2018).

Strengths

Interesting novel concept and quality results. This was a pleasure to read and will be a useful contribution to the research community.

Weaknesses

I'd be interested if you could discuss some implications or interpretations of your work. You obtain PAC bounds for the various settings and you summarize your results in Table 1. I am not immediately sure what the quality of these results are, how they compare to prior work, and how they would be represented in real-world learning. Minor: - Perhaps include a short primer on VC dimension, pseudo-dimension, Rademacher complexity, and PAC learning for unfamiliar audiences in the appendix - Perhaps cite (Xia, AAMAS 2013) and related papers on preference/rank learning (e.g., (Zhao, Liu, and Xia, IJCAI 2022) or (Newman, Royal Society 2022) or (Conitzer and Sandholm, UAI 2005) or (Xia, Conitzer, and Lang, AAMAS 2010)) as related work - Line 142: "welfare" not "malfare"

Questions

NA

Rating

7

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

Yes

Reviewer W5KA4/10 · confidence 1/52024-07-06

Summary

This paper studies the learnability of social welfare functions -- which are functions over the utility of a group of voters and an outcome. Under varying information schemes, they address the question of how well it is possible to learn the social welfare function being used by a decision maker. The first setting considers learning when cardinal values of actions are knowable, which the authors point out corresponds to regression. The second setting looks at when the information given is a pair of utilitty vectors and some indication of which vector corresponds to higher social welfare. Finally, the authors consider the pairwise model when information is noisy. The paper shows that in all settings being considered, a large class of social welfare functions are learnable with a polynomial number of samples. Experiments demonstrate the existence of a practical algorithm for these results.

Strengths

The problem being studied in this paper is well-defined and seems like it may be interesting. Despite the questions asked not being obscure, I am aware of little work that studies similar ideas ([17] being the exception, and I've always thought it odd that more work has not directly built atop [17]). The methodology taken in the paper seems quite reasonable. While the proofs are not included in the body the results appear correct, to my limited understanding. This problem is certainly able to inspire potential future research and provides a reasonable contribution in its own right.

Weaknesses

While only incidentally a weakness of this particular paper, the state of science would be better if this were three separate papers (or a journal submission). There is simply not enough time to consider the paper in depth and the parts of the paper important for peer-review (ie. the proofs and many validation experiments) are not in the part of the paper that gets the bulk of reviewing attention so my understanding of the paper is quite limited; I found the math quite dense and open to improved clarity. As far as clarity is concerned, the paper is moderately readable but I feel that the results could be explained somewhat more clearly without requiring additional space. The introduction does an adequate job of outlining the common idea of social welfare and what problem the paper studies. I understand what the problem solves but not until the end of Section 7 is there some suggestion of why this problem might be interesting. Motivating the questions in the paper earlier would be useful.

Questions

N/A

Rating

4

Confidence

1

Soundness

3

Presentation

2

Contribution

2

Limitations

Some discussion has been included but it is somewhat limited and surface level.

Reviewer 7Nn55/10 · confidence 4/52024-07-15

Summary

This work studies learning the social welfare function from a power mean function class. - They first consider the cardinal social welfare setting, where the data distribution is over the utilities and social welfare values. They provide the upper bounds on the pseudo-dimensions of the function class and then apply them to bound the generalization loss. - They then consider pairwise preference setting, where the data distribution is over utilities of a pair of actions and the comparison of their social welfare values. They provide bounds on the VC dimension of the function class and then apply them to bound the generalization loss. They further study two noisy settings, iid noise and logistic noise. - Finally, they conduct experiments to justify their results.

Strengths

This work studies a new problem of learning social welfare function. The writing is very clear. They provide both theoretical and empirical analyses.

Weaknesses

The availability of labels in the real world: In the cardinal setting, the label of the data point is the true social welfare value. I am curious if there really exists any such labeled data set. I have the same question in the pairwise comparison setting. The technical contribution seems to be limited. It looks like the results are derived by applying standard learning theory results. If there are any technical challenges in deriving the results, e.g., Lemmas 3.1 and 4.1, it would be great if the authors could address them. This is also not a new approach but standard ERM. So far, the problem studied in this work looks like a special case of general learning problems and doesn't require any new techniques. The experimental results for this problem also look consistent with the phenomena in general learning, e.g., loss decreases with decreasing noise and increasing sample size. What is the takeaway information from the experiments?

Questions

see weaknesses.

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

2

Limitations

NA

Reviewer 5Qbq7/10 · confidence 2/52024-07-24

Summary

The paper is about learning social welfare function which belongs to a well-studied family of weighted power mean function, as a way to understand a policy maker's rationale. In particular, the paper focuses on two settings: 1) when the input is the vector of utilities/social welfare, and 2) when the input is a pairwise comparison. The paper derives theoretical bounds for different social welfare information for different kinds of loss function with both known and unknown weights, and

Strengths

The paper focuses on an interesting and important question on social welfare function learning, which has great potentials in social learning and policy making. Overall, it is well-organized and well-presented. The theoretical results are solid and elegant. Overall, the authors are careful and transparent about evaluating both the strengths and weaknesses of their work. The claims/arguments are explored in sufficient depth.

Weaknesses

I find some of the results hard to interpret, for example, the bounds in Theorem 3.2. It would be great if the authors could add intuitions behind the result and better demonstrate the influence of each term.

Questions

1. How should I understand the "pseudo-dimensions" of $M_{w, d}$ and $M_d$ in line 161? 2. For the pairwise comparison setting, does it require access to the pairwise comparisons between all pairs of actions? Can it be extended to a setting where there's only partial comparison available, if so, how does it change the theoretical bound?

Rating

7

Confidence

2

Soundness

3

Presentation

3

Contribution

3

Limitations

See above.

Reviewer W5KA2024-08-12

TL;DR: The math/theory is not explained clearly enough for someone not already expert enough to write this paper themselves. Our difference of opinion is largely philosophical, and partially from my preference to write a complete review (rather than leaving sections blank). When I write, for example, that the methodology "seems quite reasonable" it is only partially a strength. My other meaning is that something has gotten lost between your keyboard (writing the paper) and my keyboard (writing the review) that makes me lack the confidence to write a stronger sentence. I suspect that our philosophical disagreement may lie in whether it is the author's job to write a paper as clearly as possible or the readers job to be smart/hardworking enough to understand it. In writing my review I went back to Ariel's 2009 paper to try to answer the question "is it possible to write something with a similar amount of depth in a clear and understandable manner?" I found that paper much more readable which suggests to me that this paper could (and, therefore, should) be more readable as well. If I were to run into this paper as a reviewer again, I would give a higher score if concepts and results were explained in a more clear manner but, for now, I maintain my current score. To be clear for meta-reviewers deciding what factors to prioritize: My score is largely based on the clarity of the paper which has prevented me from reviewing the paper in sufficient depth needed for a higher score. Not as a result of any known technical issues.

Authorsrebuttal2024-08-12

Thank you for your response. We would like to solicit any specific changes/additions you can suggest to improve the clarity of the paper, and would be happy to incorporate these. Such feedback would be very valuable to us, since we strongly believe in the interdisciplinary value of our work, and are committed to making it accessible to a broader audience. Based on the other reviewers' suggestions, we believe the following changes would make the paper easier to read: - Additional clarifications for all theorems and lemmas: While it would be difficult to incorporate proofs within the NeurIPS page limit, we can definitely add more intuition for our results. We point to our rebuttal to reviewer 5Qbq as an example, where we give intuitions behind Theorem 3.2. We will rewrite our current exposition to include these additional points. - Improved motivation for experiments: Our rebuttal to reviewer 7Nn5 contains further motivation for our experiments, and we plan to add it to the paper to better contextualize our experiment plan. - As reviewer 1QjM has suggested, we will add a primer on key learning-theoretic concepts like VC dimension, Rademacher complexity, and pseudo-dimension. Finally, while we appreciate the comparison with Procaccia et al. (2009), we note that it’s a journal paper that enjoys unlimited space, so it may set an impossibly high bar for a 9-page conference paper.

Reviewer W5KA2024-08-13

Yes, those sorts of things are exactly what I would hope to see in a very well written paper. In my opinion, a paper is most accessible if it explains ideas in multiple ways (e.g. providing an intuitive outline for a first reading/a non-expert and providing theory/proofs for deeper reading/the more mathematically inclined). Your final note touches on something I mentioned in my initial review. It may be that fitting all of this content into one paper while explaining it clearly is simply not possible. I agree that it may not be possible in the confines of a conference paper, hence my comment about multiple papers or a journal paper. My rejection suggestion should certainly not be taken as an indictment of the underlying work, but rather as (i) more explanation being needed (as discussed), and (ii) the current paper simply not being a good fit for conference publication, in my opinion.

Reviewer 5Qbq2024-08-12

I've carefully read the rebuttal by the authors, my evaluation of the paper remains the same.

Reviewer 7Nn52024-08-13

Thanks for the response! My concerns are addressed and I'm happy to increase my rating. I suggest the authors to include the discussion for the first two questions in the updated version.

Program Chairsdecision2024-09-25

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC