Summary
This paper studies a model of content creation and consumption on arbitrary online user-generated content platforms (e.g., YouTube, TikTok). It focuses on a type of Cournot competition in which creators mainly modify their creation volume. The paper provides a description of this model, a theoretical analyses of the Pure Nash Equilibria in this setting, an analysis of how platform designers might use mechanism design to balance consumer and creator utility, a framing of this balancing problem as an optimization problem solvable via (approximated) gradient descent, and experiments using purely synthetic data (sampled "users" with Gaussian preferences) and empirical data (users with preferences from the MovieLens dataset, popular in recommender systems).
Strengths
Overall, this paper provides a strong overall contribution and number of results and insights that will be of interest to a number of different communities -- researchers interested in UGC and online communities, mechanism design, ML for social media, etc.
The clarity is high throughout. The paper begins with strong and well argued motivation, the organization is helpful, and in general the overall narrative of the paper is clear.
In terms of novelty, this paper directly builds on a previous modelling work, but is very upfront about highlighting what the main differences and additions are in terms of contribution. The experiments seem to especially build off the design of [40] (esp. in terms of the synthetic data + MovieLens combination), which might be worth mentioning if that is intentional.
Overall, the potential significance of this work seems potentially high.
Weaknesses
Overall, I expect readers won't have any major concerns with the theoretical results or experiments (see some minor questions below in the Questions section).
Rather, the main threat to the significance of this paper is making the case that that a Cournot-style is actually common in the UGC platforms being invoked here. Of course, even if only a few platforms really end up being well-described by the model, the contribution is still very meaningful. That said, a few specific concerns with the current draft:
- a number of specific platforms are mentioned by name: YouTube, TikTok, Netflix, Spotify, and MovieLens.
- Only data from MovieLens is used (which is very reasonable -- it's a very popular dataset for academic work for good reason).
- However, the named platforms vary quite a bit in terms of their actual creator competition, i.e. one would expect the incentives of a platform like Netflix (which also acts a creator agent, sometimes with substantially higher budget than other creators) to differ quite a bit from TikTok
See "Questions" section below for some specific questions about this concern that I think are likely to be in scope of a revision.
With this critique in mind -- that certain platforms might violate the assumptions needed for the model to work well -- I think the current draft may overstate the generality of the conceptual insight.
Questions
A few very specific questions about the model (with the caveat that of course anything along the lines of using empirical data from major platforms and/or trying to frame this model as predictive for an entire spectrum of platform types is probably out of scope)
- What is the strongest evidence that any of these major platforms follow Cournot competition like dynamics?
- What is the impact of platform-as-creator dynamics, such on Netflix?
- More generally, it would be helpful to explicitly state how resource heterogeneity amongst creators or budget heterogeneity amongst consumers may or may not cause issues for the use of this model.
- To what extent would we expect results to hold if we did have access to MovieLens-style observational data from e.g. YouTube?
Overall, these are not "existential" questions per se, but some attempt to clarify could strengthen the draft quite a bit.
Limitations
I do think the current draft could do more to justify the strength of the conceptual claims and/or hold a bit more space to explicitly discuss limitations (primarily, how well requisite assumptions hold across the platforms of interest). See above (Questions).