Collaborative learning techniques have significantly advanced in recent years, enabling private model training across multiple organizations. Despite this opportunity, firms face a dilemma when considering data sharing with competitors -- while collaboration can improve a company's machine learning model, it may also benefit competitors and hence reduce profits. In this work, we introduce a general framework for analyzing this data-sharing trade-off. The framework consists of three components, representing the firms' production decisions, the effect of additional data on model quality, and the data-sharing negotiation process, respectively. We then study an instantiation of the framework, based on a conventional market model from economic theory, to identify key factors that affect collaboration incentives. Our findings indicate a profound impact of market conditions on the data-sharing incentives. In particular, we find that reduced competition, in terms of the similarities between the firms' products, and harder learning tasks foster collaboration.
Paper
Similar papers
Peer review
Summary
This paper uses algorithmic game theory to study potential data sharing between competing firms that train machine learning models with similar goals. The authors propose a new framework that can be used to analyze trade-offs faced by data-using agents who make decisions about whether to collaborate (via data sharing) to improve their own ML performance, while potentially losing out on profits because they also improved their competitors ML performance. The authors draw on conventional market models from academic economics to make predictions about how different structural factors might impact collaboration decisions. This analysis has implications for the development of regulations and norms pertaining to data sharing.
Strengths
This paper has a number of strengths. The motivation, combination of theoretical arguments with simulation evidence, and reasonable use of assumptions all stood out. First, the paper’s motivation is quite strong. The authors argue that the incentives underlying data sharing between machine learning operating firms is understudied, and that using the lens of theoretical modeling can highlight how market conditions might impact sharing behaviors. The work combines theory (drawing on published economics literature) and a simulation experiment. I thought this combination was convincing: the experimental component is likely to help readers understand the implications of the theoretical framework. There are a lot of assumptions at play, but they seem fairly plausible on the whole. Thinking in terms of whether this theory-focused paper could guide practical decision-making (by firms) or policy making (by regulators), the current draft devotes enough attention to arguing for plausibility. I do think there’s room to see more discussion of where these assumptions are more or less plausible (see below), though this may be out of scope for a theory paper that’s trying to propose a new framework and ask for some re-thinking.
Weaknesses
In my view, the main threat to the validity and impact of the work is the reliance on assumptions from conventional economics. The authors are very upfront about signposting which prior works are most foundational (e.g. the 1979 work defining representative consumer with quasi-linear quadratic utility, and classical models of competition that focus on quantities vs. prices) but empirical validation could help this line of work have more impact in the long-term. Put another way, the current draft does a great job of pointing out relevant literature that it builds off, but could do more to explicitly state why certain assumptions fit a specific empirical context. Of course, the authors may not wish to zoom in on a specific context, which is a fair choice for scoping the paper. The paper also discusses data as a “key asset” (e.g. in the Introduction) in a very broad sense. None of the claims are situated relative to specific use cases of data and ML. Examples like ad tech and a “production process” are gestured at, and of course the running example of taxi driver scheduling is helpful. On the whole, however, I think it may be hard for readers to assess how contextually dependent some of the claims and results are. Overall, the impact of the work could be improved by clarifying the extent to which these analyses are or are not contextually dependent (even in terms of specific numerical characterizations of data scaling behavior / “data impact model”).
Questions
Early in the Introduction, free-riding and non-collaboration concerns are mentioned. The authors may want to engage briefly with sociological work on collective action (though perhaps this better suited for future work and out of scope for this paper). I think a big open question for future work along these lines is which types of data are best handled with an “economic-leaning” model vs a “sociological-leaning” model that includes non-monetary incentives facing individual data generating agents. In general, the framework also seems to lack any notion of the possibility of commons or public goods. I could imagine adding an additional stage (or several) to the game to account for this. Fully engaging with this topic is almost certainly out of scope given the current space constraints, but perhaps worth a brief mention. While I expect the core audience for this kind of paper (i.e. readers familiar with some of the references already or interested in data collection games) will follow most sections, there’s potential to strengthen the broader impact of the paper just by adding a bit more high-level summarization of each sections.. Sections 3.1 and the end of Section 5 do a nice job of this kind of discussion, but there’s opportunities to emphasize this kind of recap in other sections. As a minor comment, it may help the paper to discuss whether data-dependent products lend themselves to Bertrand vs. Cournot competition. It seems Cournot may be preferred, but I didn’t quite understand why. This relates to my main high-level comment about the work, which is that I think readers will want to know how different assumptions here map to different data-dependent technology contexts.
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
4 excellent
Presentation
3 good
Contribution
3 good
Limitations
The work is reasonable in discussing its own limitations. In terms of societal implications, I think the discussion in the paper is commensurate to potential concerns. Overall, the paper is a bit light on discussion of how data sharing would work in practice (not touching on topics like public goods, anti-trust concerns, consumer welfare, etc.) but I think this is reasonable given space constraints and paper scope. As noted above, I think some engagement with literature on collective action, commons, and public goods could be insightful, or even a brief mention of why the authors think the collective action / commons perspective on data-dependent technologies is not relevant.
Summary
The authors present a trainable model for optimising the benefits of data sharing based valuation of data, benefits of the shared data sets for the inference model, in a conventional market model. The authors point out that the model can be used to find an optimal data sharing strategy. The model is nicely developed and mathematically sound, presented clearly in an understandable manner. It states the assumptions that are used in the development of the model and provides a good a good summary of the consequences, the effect of the simplicity of the learning task, and similarity of the products build using the shared data sets. The mean coalition size is shown as function of these parameters.
Strengths
The paper is written and goes through the relevant steps in creation of the model. It provides excellent argumentation and conclusions from the assumptions made.
Weaknesses
The biggest weakness I see in the paper is its assumption that the collaborative partners share the same data distribution and the data is modelled as iid samples of a common distribution. This is generally the situation when there less sense of sharing data, as one has the capability, in time, to the get a representative data set that is enough for an highly accurate inferred model. However, the real problem, is in a case that the data sharing partners work with differently distributed, where the parties are do not have the capability to reproduce the data that the other parties have. This is for example the case where medical X-rays are taken with different X-ray machines, in diverse set of hospitals and the aim is to build an inference model that works generally with a variable types of machines. An other example is, again in the medical domain, where the ethnic origin of the subjects has to be taken into account. Then one cannot really build a good model that is not ethnically non-discriminating without schemes of sharing data. An other aspect that has not been taken account, is the effects of legislation, for example EZU Data governance act that mandates data sharing for data recorder from what the act calls connected devices. Also, the effects of data sharing based on EU GDP rights for subjects, not companies, to share their personal data with third parties. These require more complicated data strategy consideration that should be discussed to make the paper valuable for real evaluation of companies data sharing sharing strategies
Questions
I would like the authors to address the points discussed in the weaknesses part of the review, especially redoing their data value analysis based on cases where not including the data from others would lead to discriminative AI models that are generally not acceptable. Also, even the current analysis should contain the estimation how long time would it take to build an own model (and estimate the cost of being late in the market with AI features) compared to the complications of sharing the data. The equivalent extra time for a attaining a similar product alone could be estimated already with the current model in the manuscript.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Soundness
3 good
Presentation
4 excellent
Contribution
2 fair
Limitations
The libations are discussed in the weakness part - as these are the major weakness, of the otherwise very good paper.
Summary
This paper explores the dilemma faced by firms when considering strategic data sharing with their competitors. It introduces a framework to investigate the incentives for data sharing and examines its impact on collaboration and profitability. The author discusses the barriers to data sharing, such as privacy concerns, and proposes a market model, data impact model, and collaboration scheme as components of the framework. The findings suggest that reduced competition and harder learning tasks foster collaboration. An illustrative example of a taxi market is provided to demonstrate the concepts discussed. Overall, this study aims to understand how market competition affects collaborative learning incentives and provides insights into the data-sharing trade-off.
Strengths
- The study introduces a novel and comprehensive framework to analyze data sharing between competitors, considering factors such as machine learning model quality's impact on production cost. - The study investigates the incentives for data sharing and examines the impact of market conditions, product similarities, the complexity of learning task and firm size on collaboration incentives. Albeit it's simply based on a conventional market model and a natural model of data impact grounded in learning theory, the findings are inspiring and the exposition of the results is clear and easy to follow.
Weaknesses
- Since this research is primarily about economic modeling and analysis, it would be great to see a discussion of the various real-world examples in the context rather than just one taxi market example and a oil-market example in the appendix. - This study uses simulation to examine collaboration incentives among multiple firms, but it would be valuable to discuss the limitations of the simulations and potential biases inherent in such studies.
Questions
- The study mentions the use of a data impact model grounded in learning theory, but it doesn't provide detailed information about the model's validity and discussion about alternative formulations. Is there any other forms of data impact model could also be considered? - I can see there are plenty extensions provided in the supplementary materials, is there a way to get comparable results for an analog of the problem that consider a sequential game(such as Stackelberg) rather than simultaneous game? This set up may happens in some mega tech companies get the drop on some other small firms and most often the big companies take the advantages. - The big companies tend to have more bargaining power than small size firms, and the cost functions are often asymmetric. Does Theorem 6.1 still holds when bigger company has different cost function as it is discussed in C.3?
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
The authors didn't discuss the limitations of their theoretical results. This study is mainly theoretical and primarily focused on introducing a novel framework on strategic data sharing dilemma. There are still plenty rooms for in-depth investigations into the data sharing studies.
Summary
* This paper analyzes the economic consequences of data sharing between competitors. * Interaction is modeled as a market: * Market has $m$ firms ($F_1,\dots,F_m$), where firm $i$ produces $q_i$ units of good $G_i$ at price $p_i\ge 0$, and quality $v_i$. * In addition, there are “outside goods” $\{1,\dots,k\}$ offered at fixed prices $\tilde{p}_l$. * Each consumer $j$ optimizes their utility $u^j$ by deciding on consumption of firm goods $q^j\in\mathbb{R}^m_+ $ and outside goods $g^j\in\mathbb{R}_+^k$ under budget constraint $B^j$ - Leading to Firms maximize their expected utility $\mathbb{E}_v[p_i q_i - C_i(q_i,v_i)]$. * Different solution concepts are considered: Firms act either by deciding on $q_i$ (Cournot competition), prices $p_i$ (Bertrand competition), or by strategically considering the response of their competitors (Nash equilibrium). * In section 4, a concrete market model is instantiated: Consumers make decisions according to a quasi-quadratic utility model characterized by a substitutability parameter $\gamma$, and the cost associated with each firm is $C_i(q_i,v_i)=c_i q_i$, where $c_i$ depends on the quality of the machine learning model. Machine learning quality is assumed to affect production costs only. * For analysis, data is assumed to be homogeneous (data of all firms is sampled independently from the same distribution), and coefficient $c_i$ is assumed to take a concrete power-law parametric form ($c_i = a+b/n^\beta$). Attention is restricted to competition between two firms. * Theorem 5.1 characterizes the equilibria as a function of the $\gamma$ parameter, and the ratios between the amounts of data collected by the two firms. Theorem 6.1 characterizes the equilibrium in the case of partial data sharing. Finally, a numerical simulation is conducted on a competition setting with more than two firms.
Strengths
* Topic is well motivated. * Clean presentation, connects contemporary topics in machine learning to classic economic notions in a creative and interesting way.
Weaknesses
* Parametric assumptions for main theorems are not validated against real-world data. Not clear which parameter regimes are likely in practice. * Data homogeneity assumption may be too conservative - Homogeneity means that datasets collected by all firms are assumed to be sampled independently from the same distribution. This is unlikely to be the case - For instance, in the taxi running example, different operators are likely to observe different data distribution, e.g. because they operate in different parts of town. * Limitations of the method are not thoroughly discussed.
Questions
* In the proposed model, do the firms have the ability to invest part of their budget in independent data collection? (e.g by conducting surveys to improve training set quality, or buying data from an external provider which is not a direct competitor). If the market model does not include this possibility, what would be the consequences of considering it? * What are typical values of the constants $a$,$b$,$\gamma$ in real-world systems, and why? * L224: “We assume that $b/(1-a)$ is small enough...” - Which values of $b/(1-a)$ are small enough for the lemmas to hold? Are they realistic in real-world systems? * How would the results be different if the cost $c_i$ had a different parametric form? For example, using similar arguments to the ones presented in the paper, one could consider costs that relate to the parametric forms described in Viering et al. 2022 (Table 1).
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
Limitations are not thoroughly discussed. In particular, I feel it would be helpful to provide an indication of where we expect the main assumptions in this paper to hold, and discuss the consequences of making wrong assumptions.
Rebuttal helps address areas for improvement
Thanks for this rebuttal, authors! Overall, I think the points raised here are fair (e.g., challenges with incorporating public goods directly into this current work, the discussion of using global welfare lens). I thought the general argument for avoid hyper-specific data use-case here is fair as well. Overall, I appreciate this additional information from the authors, and it sounds like it will be possible to build on some of the areas for improvement in camera ready. I think this paper will be a valuable contribution to the conference.
Thank you for your response!
Thank you for your timely response! We appreciate your constructive and positive feedback, which we will incorporate in the next version of the paper. In particular, we will discuss the non-rivalry of data and the obstacles in front of collective action in collaborative learning, as well as the positioning of this work as a general framework that opens the door for application-specific models.
Thank you for the clarifications The consent coefficient is a good addition. It also covers the case of copyrighted data, as there is push (despite of the text and data mining exception) to ask licences for training use of copyrighted works, and companies start to voluntarily accept this. On the training cost case use case, I really meant foundation models, where it really matters. For training these, gaining as large, and diverse data as possible is crucial. I think the considerations of this paper are of greatest value in this context, as most of the advances in AI, are currently emerging from transformer like technologies, be in protein folding, automatic coding or image generation etc... The need for large models will overshadow others, not only in use, but also in costs of data.( and in human labelling based grounding), in costs of training, and costs of running. There is an entanglement of these in the deployment that one cannot factorise easily into independent components - like stand alone data strategy, without interfacing the others factors. In this context, there is also other "data" that can be shared and is valuable, like encodings of texts (and images), and the weights of the pre-trained neural networks that can reduce training costs, also it can be reduced by giving out data and asking the recipient to train the model, and get back the trained model. This way of paying training costs with data is actually quite common. Training costs can be estimated by quoting AWS, Azure, Google cloud,... GPU time pricing.
Thank you for your response!
Thank you for your timely and detailed response! We completely agree with the reviewer that foundation models are increasingly prevalent, and their training costs are not negligible. We will be happy to provide a discussion in our manuscript regarding training costs as a possible consideration that may additionally come into play in the context of foundation models. While training costs are certainly interesting to model, we note that they give rise to many orthogonal incentives to the trade-off studied in this work. In particular, sharing computation brings up the aspect of fair client and server compensation for training costs, which is a different line of work in the FL incentives literature (Tu et al., 2022). Additionally, we see several obstacles in front of the direct modelling of training costs in the context of our problem. Specifically, this will likely require application-specific (and potentially proprietary) information on how these costs are actually incurred. First, it is unclear how the companies will negotiate the training costs splitting. For example, they might split the costs equally among all coalition members, they might split them proportionally to the sizes of their datasets, or, as the reviewer suggested, the central server might bear all the costs, but receive data as compensation. Similarly, there are multiple options how the firms will do inference. They may receive a copy of the end model and use it locally, or they may leave the model at the central server only. If the latter case, it is unclear how they will pay for inference. For instance, they might pay for each query, they might get a quota proportional to their data contribution, or they might auction the server inference time. We will be happy to elaborate on these considerations in the next version of the manuscript. Tu, X., Zhu, K., Luong, N. C., Niyato, D., Zhang, Y., and Li, J. Incentive mechanisms for federated learning: From economic and game theoretic perspective. IEEE Transactions on Cognitive Communications and Networking, 2022.
Thank you for the thorough and helpful response! In particular, I appreciate the order-of-magnitude estimation of model parameters, and I believe that such grounding significantly strengthens the presented results. Given the suggestions you made for improvement, I view the paper as a step in the right direction, and I believe it can facilitate fruitful discussions within the community. I'm increasing my rating to 7 (Accept).
Thank you for your response!
Thank you for your timely response! We appreciate your constructive and positive feedback, which we will incorporate in the next version of the manuscript. In particular, we will include the discussion on the model parameters $a, b, \gamma$.
Thank you for the response. I have a followup comment to Q2 and Q3: - The setup I'm interested in is when some mega tech companies who has different form of costs(maybe lower in \beta_i since they got more resources) than small tech companies and oftenly move first in the market(a Stackelberg game). This may be an overly-detailed setup, but it's more common in real life. And it's intuitively a potential special case for scenarios when companies find it hard to collaborate. It may provide more insights to whole context if it's turned out to be true.
Analysis of a Stackelberg setup
Thank you for your constructive feedback and for the clarification! Following your suggestion, we investigated a setup where the competition phase corresponds to a Stackelberg game between two companies. As you suggested, the big company $F_1$ has a better cost function ($\beta_1 > \beta_2$) and more data ($n_1 > n_2$). We consider the full data-sharing negotiation scheme from Section 5. Repeating our two-firm analysis (see Appendices A.2, A.3, and A.4) for the Cournot-based Stackelberg game (e.g., Boyer & Moreaux, 1987), where the first firm decides on quantities before the second one, we get the following collaboration criteria: $$\varPi_{1, \text{ind}}^e \le \varPi_{1, \text{share}}^e \iff \gamma (n_2^{-\beta_2} - (n_1 + n_2)^{-\beta_2}) \le 2 (n_1^{-\beta_1} - (n_1 + n_2)^{-\beta_1}),$$ $$\varPi_{2, \text{ind}}^e \le \varPi_{2, \text{share}}^e \iff \gamma (n_1^{-\beta_1} - (n_1 + n_2)^{-\beta_1}) \le \Bigl(2 - \frac{\gamma^2}{2} \Bigr) (n_2^{-\beta_2} - (n_1 + n_2)^{-\beta_2}).$$ Here, the first company's incentives to collaborate do not change compared to the Cournot case, while the second company's incentives decrease $\Bigl(2 - \frac{\gamma^2}{2} - \gamma \le 2 - \gamma \Bigr)$. Despite the reduced incentives for the second firm, since $\beta_1 > \beta_2$ and $n_1 > n_2$, the second condition always holds, both in the Stackelberg setup presented here and in the context of Theorem C.1 in Appendix C.3. Therefore, the smaller company will always want to collaborate. Since the incentives of the first firm are unchanged in both cases, there is no change in the likelihood of collaboration compared to the Cournot case. Interestingly, the first firm will have larger profits in the Stackelberg setup, compared to the setup in Appendix C.3, since it could always choose Cournot equilibrium quantities at the first stage of the competition and get the Cournot equilibrium (Anderson & Engers, 1992). We hope the reviewer finds this analysis interesting and relevant to their proposed setup. Please let us know if you have any further questions; we will be happy to address them. Boyer, M., & Moreaux, M. On Stackelberg equilibria with differentiated products: The critical role of the strategy space. The Journal of Industrial Economics, 217-230, 1987. Anderson, S. P., & Engers, M. Stackelberg versus Cournot oligopoly equilibrium. International Journal of Industrial Organization, 10(1), 127-135, 1992.
Decision
Accept (poster)