Monoculture in Matching Markets

Algorithmic monoculture arises when many decision-makers rely on the same algorithm to evaluate applicants. An emerging body of work investigates possible harms of this kind of homogeneity, but has been limited by the challenge of incorporating market effects in which the preferences and behavior of many applicants and decision-makers jointly interact to determine outcomes. Addressing this challenge, we introduce a tractable theoretical model of algorithmic monoculture in a two-sided matching market with many participants. We use the model to analyze outcomes under monoculture (when decision-makers all evaluate applicants using a common algorithm) and under polyculture (when decision-makers evaluate applicants independently). All else equal, monoculture (1) selects less-preferred applicants when noise is well-behaved, (2) matches more applicants to their top choice, though individual applicants may be worse off depending on their value to decision-makers and risk tolerance, and (3) is more robust to disparities in the number of applications submitted.

Paper

References (47)

Scroll for more · 35 remaining

Similar papers

Peer review

Reviewer LFB27/10 · confidence 4/52024-06-23

Summary

The authors are focused on matching markets in which different firms use a single algorithm / evaluation criterion (monoculture) vs. markets where different firms may each have different evaluation algorithms / criterion (polyculture). This can be seen as a substantial generalization of the wonderful work of Kleinberg and Raghavan [35] on monoculture in hiring with two firms. The authors first introduce the continuum matching market model introduced by Azevedo and Leshno [10]. Here, there is a continuum of applicants, and a finite number of firms. The authors make the assumptions that (1) each firm has the same fixed capacity, and (2) not all applicants will eventually be matched with a firm. The authors first review critical results in the existing continuum model. These include the fact that a stable matching corresponds to a particular cutoff vector, and subject to this cutoff vector, applicants always choose their highest preferred firm for which their estimated quality is higher than the cutoff for that firm. Next, the authors introduce their notion of mono and polyculture into this model. Intuitively, monoculture is where each firm has an identical estimate of the value of an applicant of type $\theta(v)$, given by $v + X$ for an $X$ drawn from some noise distribution $D$. This captures, for example, each firm using Chat GPT to evaluate the resumes of all applicants. In polyculture, each firm $i$ may have a different estimate $v + X_i$ for the value of applicants of type $\theta(v)$. The authors begin by proving that the cutoff characterization of stable matching is unique in the mono and polyculture settings (Lemma 2). This follows from the lattice structure of stable matchings. Then, in Proposition 3, they show that the probability of an applicant of type $\theta(v)$ being matched under polyculture is related to the maximum of $X_i$ over all firms' noisy estimates $X_i$, whereas under monoculture this probability is related only to $X$ (since all firms have identical estimates). We now move to the main results. In Theorem 1, the authors show that under polyculture, as the number of firms $m \to \infty$, the (firm-) optimal welfare can be achieved by the resulting matching. This does not hold for monoculture. In particular, under monoculture, the probability that an individual of type $\theta(v)$ is matched at all is constant for varying $m$. In Theorem 2, the authors examine applicant welfare. They show that applicants have a higher chance of being matched with their top choice under monoculture, but that for a subset of applicants of positive measure, the variance in whether they are matched or not is higher under monoculture than polyculture. This means that not all applicants are incentivized to prefer monoculture unconditionally. Finally, some extensions under a differential application access setting are provided. Intuitively, the authors show that more applications do not help applicants under monoculture but does under polyculture. Experiments complement most of the theoretical results, and also demonstrate that the uniform preference assumption is not essential to practical relevance of the results.

Strengths

The paper is generally extremely well written, motivated, and clear. I also think that the work is already very important in the modern context in which universities and hiring managers may already be using one of only a handful of services to conduct automated applicant filtering. The work examines what this would potentially lead to in terms of macroeconomic market dynamics. The theoretical results are presented clearly, and I understood most even though I have not personally worked in the continuum matching model (have only worked in the discrete matching model). I appreciate that the authors also empirically investigate the (strong) assumption that all applicants' preferences over firms are drawn uniformly at random. The empirical results confirm that this is perhaps not a fundamental assumption, even though the (current) proofs critically hinge on it. This work certainly challenged my preconceived notion (“monoculture=bad”) in a fundamental way and may open a more general line of inquiry into monoculture more broadly. This paper was a pleasure to read, and I look forward to additional work from the authors.

Weaknesses

Note that I did not carefully check the proofs. (W1) I think the introduction of the continuum model could be made a bit more clear, in particular the definition of applicant types. For example, there seems to be a small typo in lines 141-142:: “The realization of θ(v) is their type, which lies in $\Theta \coloneqq \mathcal{R} \times \mathbb{R}^m$ is the set of applicant types,”. Further, I am not sure if this is a typo as well (in line 143): “$\succ^\theta$ is the preference ordering of v over firms…”. Do all applicants of value $v$ have the same type $\theta(v)$? That is, do all have the same preferences over firms? Or, do we draw different preferences uniformly at random for each individual of value $v$? These were not clear from just this introduction on the continuum model. (W2) The model considers identical noise across all applicant “types”. This is certainly a reasonable form of polyculture to analyze, however, in practice we may be more concerned with bias based on different “types” of individuals. I.e., historically underrepresented minorities having a skewed or higher variance noise distribution. “Types” in this sense is (I believe) not captured by solely preference and firm quality estimates. This is certainly less of a weakness and more of a direction of future work, but I think it is perhaps important to mention. The authors have some discussion in lines 350-352, but more could certainly be added earlier in the paper. (W3) I think that the assumptions made throughout the work are sprinkled throughout the paper. Having a collection of all assumptions, perhaps in the appendix, may help the reader better understand the limitations of the work. Minor: Most non-theorems from the main paper are referred to incorrectly in the appendix, e.g. Lemma 2 is mistakenly referred to as proposition 2 in the appendix, and similarly proposition 10 / lemma 10, and corollary 4 / proposition 4.

Questions

Q1: How does Lemma 2 (Equal Cutoffs Lemma) relate to Theorem 1 part 1 from Azevedo and Leshno [10], which says that if $\eta$ has full support, then there is a unique stable matching? Q2: Do we expect the maximum concentrating distribution assumption to hold in, e.g., the included experiments?

Rating

7

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

I do not believe that the authors have a dedicated limitations section of their paper, which I recommend. Some limitations are sprinkled throughout the paper (e.g. Line 350-353), but perhaps more could be mentioned. In particular, how strong/important the technical assumptions are in practice are could be explicitly discussed in a formal limitation section.

Reviewer FF9T5/10 · confidence 3/52024-07-11

Summary

This paper studies the monoculture problem in matching markets from a theoretical perspective. The authors found that on the firms' (colleges') side, monoculture may decrease the quantity of matched applicants; while on the applicants' side, monoculture may help matched applicants to match with higher-ranked firms (colleges). Additionally, monoculture may decrease the risk of unfairness if some applicants are born with more opportunities to apply to more firms (colleges).

Strengths

1. The paper studies the monoculture problem with multiple (possibly infinite) homogeneous firms and heterogeneous applicants differed by real-numbered type, broadening the elements modeling within the monoculture literature. 2. The paper provides positive theoretical and empirical results for monoculture, which are inspiring as the results are counter-intuitive and challenged the preexisting beliefs about monoculture.

Weaknesses

**General Weakness:** 1. The connection between this work and CS/ML conferences is vague, as this paper mainly addresses monoculture, which is intrinsically an economic problem. Additionally, the technical derivation seems quite straightforward. 2. In the model, firms are assumed to be homogeneous, which might be an oversimplification and not realistic. **Corrections for Typos:** 1. In line 203, it should state "the cutoff under monoculture must be lower."

Questions

See above Weakness.

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

One concern is whether the paper is a good fit for NeurIPS acceptance. The paper focuses more on economics and matching markets than on machine learning and computer science, attempting to analyze the phenomenon of monoculture using economic models. This paper might be more suitable for conferences like EC and WINE.

Reviewer dgPg3/10 · confidence 5/52024-07-15

Summary

The paper considers a matching model with a continuum of students/applicants and m colleges/firms where the firms have a noisy estimate of the candidates' quality and compares the stable matching outcome in two situations: monoculture (where all firms have the same estimate) vs polyculture (where each firm has an iid estimate). It makes two major assumptions : all firms have the same capacity and candidates' preferences are uniform amongst the m firms. Then the unique stable matching is described by a single cutoff for all firms. The paper states two main results: - Thm 1: with polyculture, as m grows large, the stable matching approaches an optimal cutoffs mechanism on the true quality; whereas with monoculture it does not - Thm 2: the probability of first choice match is higher under monoculture - Thm 3: if candidates can submit variable length preference lists, the students submitting more benefit from polyculture but not from monoculture

Strengths

The topic of the paper is clearly important. From what I see, the paper is not original except in rewording things studied in previous works as monoculture vs polyculture instead of correlation vs independence. Perhaps this can contribute to increasing the volume of literature labeled as studying the effect of monoculture (which is an important concern), but other than that I don't see anything fondamental it brings. The paper is clear. The take-aways are interesting but besides their lack of novelty their significance is diminished a lot by the very strong assumptions made (see below).

Weaknesses

There are two main weaknesses to the paper: (i) lack of novelty in the model, the questions and the flavor of the results and (ii) very strong assumptions that make the results mostly trivial so that I could not identify any strong technical result either. I elaborate on both in the remaining of this box. - The paper claims (l. 38) that the technical contribution is a matching market model that can be used to analyze monoculture. However, the model used is a particular case of that of [13] (because in [13] they can have multiple demographic groups). Even [13] with a single demographic group is more general because it can handle any level of correlation and not just 0-1 (mono-polyculture). The particular case of latent quality+noise is described in appendix A.4 of [13]. Note that in [13], some results apply only for 2 colleges, but the model applies to any number of colleges. - The paper makes two very strong assumptions: each firm has the same capacity and candidates' have uniform preferences. Under these, it shows (actually, it just states, because it is trivial) that at the stable matching each firm uses the same cutoff. This makes the rest of the paper technically straightforward, but this is very unrealistic in practice. So, even if one were to see the paper as an extension of some results of [13] for m firms, this would be only under *extremely* simplifying assumptions that make the extension straightforward and of much diminished significance. Just in contrast, most of the technical difficulty in [13] seems to be in handling the fact that the cutoffs may be different and hence if one increasing, it does not imply that the other does so too. The authors in fact show that the property of diminishing cutoffs is no longer always true for >2 firms. Of course, it holds with all equal cutoffs but as mentioned above, this is too simplifying (and trivial). - Thm 1 gives an interesting message but it is very straigthforward under the assumptions mentioned above. - Thm 2 is very similar to [13], the extension to m firms does not seem significant. See my discussion above. - Thm 3 is very connected to [7]. This is (really) discussed only in the appendix (l. 644-651), but it appears that the additional value of Thm 3 compared to [7] is minimal. - There are some numerical simulations. Unfortunately, as far as I could see, they do not relax the assumption of same capacity. Overall, even though the paper is well-written and interesting, I cannot identify any strong or novel result that would get close to the NeurIPS bar. Perhaps if the authors focus on the elements that are novel and try to remove the strong assumptions that make the results straightforward, this would improve the paper's contribution. [13] Castera et al. EC'22. I used for this review the latest available version from May 2024 https://hal.science/hal-03672270v6, but I checked and the previous version is extremely similar. [7] Arnosti. MS (2022)

Questions

See weaknesses.

Rating

3

Confidence

5

Soundness

2

Presentation

3

Contribution

1

Limitations

I very strongly encourage the authors to be much more upfront about the limitations of their model (discussed at length above). I have not seen them mentioned in abstract or introduction, whereas they are extremely important for the results to hold (in fact, some results do not hold without, see the counter-example at the end of [13] for more than 2 firms).

Reviewer V2nV6/10 · confidence 4/52024-08-05

Summary

This paper examines the effects of algorithmic monoculture in a large two-sided matching market, in which participants on both sides compete with each other and outcomes are determined by preferences on both sides. It proposes a matching markets model to study monoculture and produces both expected and surprising results. While under monoculture, all firms use a single shared estimate of an applicant's value, under polyculture, they obtain separate, independently drawn estimates. - One expected result is that polyculture benefits firm welfare. - More surprisingly, another result shows that applicants are better off under monoculture, yet risk-averse well-qualified applicants would rather prefer the security afforded by polyculture. In other words, and expectedly in this regard, monoculture presents a risk of systemic exclusion to certain more qualified applicants. - A third result is that polyculture benefits applicants who submit more applications, thus allows differences in the number of applications submitted to harm firm welfare.

Strengths

S1. Connects literature on monoculture and matching markets. S2. Produce novel results. S3. Computational experiments training on real data.

Weaknesses

W1. The results in Figure 5 seem to contradict the strong claims made about polyculture outperforming monoculture; the related claims need to be qualified. W2. The concept of "positive label", or binary 0-1 outcome, is not defined. W3. In the experiment, applicants have uniformly random preferences, which do not yield competition as suggested.

Questions

What happens when applicants have competitive preferences?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Yes.

Reviewer LFB22024-08-07

I thank the authors for the careful response. In general, I agree with the sentiment that simple proofs are in some sense desirable. Thank you for also explaining the relation (and differences) to Castera et al. I am happy to keep my score as is.

Authorsrebuttal2024-08-08

We appreciate the reply, and are glad that you found our responses helpful. Thanks again for the comments.

Reviewer FF9T2024-08-08

I'd like to thank the authors for the response regarding the contributions on techniques and novelties of the paper (though mainly within the response to Reviewer dgPg), as well as the careful clarifications about the fit-to-venue. I'm not familiar with the monoculture literature, thus I hold a conservative opinion towards the novelty contribution of this paper. I agree with the authors that the straightforward techniques are advantages in the sense of showing the insight. The model contribution seems to firstly extend the literature to the multiple-firm setting, yet weak assumption of homogeneous firms is still not a strong reason for acceptance. Although the supplemented experiments demonstrate that the paper results seem to be correct beyond the homogeneous firms assumption, it seems that for the orientation of this paper, it should be theoretical results rather than experiments that determine the acceptance. Overall, my positive evaluation of this paper remains unchanged. Besides, the meanings of $\beta$ and $\gamma$ are unclear in the attached pdf.

Authorsrebuttal2024-08-08

Thank you for the reply, and we're glad that you found our responses helpful. Thanks also for the note about $\beta$ and $\gamma$. These are defined as in lines 743-746 in the paper (in the appendix). $\beta$ controls the level of "global correlation" in applicants' preferences over colleges. $\gamma$ controls the level of "local correlation" (i.e., how much applicants prefer to be close to "nearby" firms).

Authorsrebuttal2024-08-09

We respond to each point below. Fundamentally, we believe that Castera et al. (a nice paper!!) studies a different question than us, our results/insights/techniques are conceptually new, and that another paper stating that their model can extend to ours does not equate to **findings** that affect novelty. **“First, it is not true that [13] does not model true value (which they call latent value), as the added comment of the authors seems to acknowledge.”** Again, we disagree with this comment. The model in [13] is only stated in terms of estimated preferences (which can be *motivated* by a latent value model). Consequently, and most importantly, no results imply anything about students in terms of their latent value, as all our results do. **“Also, regarding m colleges (instead of 2): the model of [13] is stated with 2 colleges, however they mention in the discussion that the model trivially extends to m colleges; so I was not able to understand the sentence stated in the rebuttal, ‘One technical insight of our work is showing how a Azevedo-Leshno continuum model admits a tractable way to study M > 2 schools’.”** While it is of course true that the model in [13] could be stated with many colleges, we believe the relevant factor is if the **results** are extended. Castera et al. does not prove many-firm results; we show that some effects *only emerge* when considering many schools (for example, comparing to results in Kleinberg and Raghavan). Moreover, increasing the number of schools *increases theoretical tractability,* especially when combined with studying maximum-concentrating noise (e.g., Gaussian, bounded). **Concerns about experimental setup: “First, random settings are often not the pathological ones that challenge the theoretical results. Second, they are also not happening in practice.”** We believe that (1) our numerical experiments capture the important/realistic types of variations in real markets, and (2) our ML-based experiments (extending those of Bommasani et al.) are more realistic than existing theory/simulations and speak to the particular community of interest. The focus of our experiments is not to identify potential “pathological” counterexamples, but rather to test our predictions in a range of settings that we believe mimic real markets. Our simulation setup focuses on two types of correlation widely noted in the literature: correlation arising from shared vertical preferences, as well as from horizontal preferences (preferences that depend on “proximity”). These are controlled by $\beta$ and $\gamma$ parameters (see lines 740-746), which we vary widely. We also simulate firm preferences generated using ML models—testing our predictions in a more realistic and relevant decision-making setting. More broadly, we think a key role of a model is to convey intuition and provide useful predictions. For example, even though Castera et al. note that “most of our results do not extend to more than two colleges” (page 5), we think that their two-college results are useful to understand broader phenomena (e.g., increasing correlation in one group can help all groups). In fact, Fig. 5 in their paper, even if it does not *exactly* match their theory, seems to suggest that the theory a very good approximation in this “pathological” example.

Reviewer V2nV2024-08-14

Your reply should mention whether there is something to be seen in the current pdf, and if so, where that is to be seen. The statement "these can be seen in our PDF" is not helpful.

Authorsrebuttal2024-08-14

PDF details

We apologize. The PDF reference regarding applicants having shared preferences over firms refers to the middle pair of plots, labeled "Texas, Correlated Applicant Prefs" and "California, Correlated Applicant Prefs". Thank you for considering our response!

Reviewer V2nV2024-08-14

The respose remains unhelpful; such plots do not exist in the provided pdf.

Authorsrebuttal2024-08-14

Plot/pdf location

Hi, I think there might be some confusion. The pdf we are referring to is the one page pdf attached to the author rebuttal to all reviewers at the top. In the middle of the pdf there are 2 sets of plots next to large text on the left that says "Texas, Correlated Applicant Prefs" and "California, Correlated Applicant Prefs", respectively. These are the plots where the x axis says "subsets of models" and y axes say "accuracy" or "average rank of match". There are many red and blue dots in these plots. These replicate Figures 3 and 5 of the initial submission, showing generally that under monoculture applicants rank to their more preferred matches, even when their preferences are correlated.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC