Rebuttal by Authors (part 2/2)
**Connections to past work.** The reviewer suggests connections to the mechanism design literature. We disagree: our proposed framework differs from mechanism design. Algorithmic mechanism design aims to maximize some social surplus (e.g., sum of utilities for all agents, in the language of our framework) by learning an optimal allocation rule (e.g., what we call a “model”) of some “resource” to agents [A3]. In contrast, our model/allocation rule is fixed, and we study how multiple individual agents respond. This aligns with our motivating problem setting in U.S. Medicare: given a fixed model (CMS-HCC) for calculating payouts to healthcare providers/insurance companies, providers/companies adjust their behavior accordingly. A mechanism design approach could inform upgrades to the CMS-HCC model, but is not our focus.
We believe that our review of game-theoretic work in machine learning is extensive, though we are happy to consider suggestions to discuss specific papers. To that end, the closest related area is strategic classification, in which agents maximize their utility function in response to an ML model’s decisions. To that end, we have provided citations to related settings in L34-L38. We pay special attention to works with similar assumptions to ours, *e.g.*:
* Utility functions are partially known (their setting) [A4] vs. utility functions are unknown beyond the general form (our setting)
* Multiple agents with different capacities to game, modeled by norm constraints on manipulation (their setting) [A5] vs. different cost-scaling factors (our setting)
We will consider a separate section/heading for the related works.
**Model is “very marginal.”** We disagree that this is a weakness: small changes to the problem formulation may significantly affect the solution space. However, in addition to a novel utility-maximization formulation for gaming in multiple agents, we contribute theoretical analyses of the resultant utility-maximization problem, which motivates a causal effect estimation approach to producing a ranking of agents by gaming propensity. To our knowledge, our approach is the first to adopt a causal effect estimation approach to gaming ranking (though previous works have explored causal mechanism design approaches to counter gaming [A6, A7]).
**Evaluation on one synthetic dataset => limited generalizability?** We partially disagree: a synthetic dataset for evaluation is necessary to establish proof-of-concept of the proposed approach. This is because ground truth gaming labels are inherently difficult/infeasible to obtain in practice (e.g., require accurate audits of all agents, or accurate self-reporting of fraud). Synthetic data validation is standard in causal inference (e.g., [A8, A9] are two well-established causal effect estimators that use IHDP [A10], a synthetic dataset, for validation).
To further mitigate generalizability issues, we verify our rankings align with known drivers of upcoding (e.g., the state-level prevalence of private healthcare providers [A11, A12], Section 4.3, Table 1, pg. 9). We also align our design choices with our problem assumptions (Appendix C.1, pg. 16-17). We are happy to consider specific suggestions for other datasets or design choices that could improve the generalizability of our findings.
**Are assumptions common in the literature? How are they justified?** We justify our assumptions after their introduction on L92-110, with examples where applicable. We like the suggestion to discuss how common our assumptions are in the literature; in particular, many strategic classification works implicitly use some form of Assumption 1 [A6, A13], and our Assumptions 2-4 are more general cases of assumptions found in [A4, A13, A14]. Assumption 5 ensures that the ground truth is predictable from the observed variables $x_i$, and is related to assumptions of no unmeasured confounding [A15]. Assumptions 6-8 are standard assumptions in the causal inference literature (e.g., Chapter 3 of [A15]). These hold independently for all agents. We will add these points to the revision.
**Other changes.** Thank you for your diligence in checking our notation — we appreciate the careful notes! We will double-check that all variables are well-defined and identify places where we’ve used similar notation/overloaded notation.