Who's Gaming the System? A Causally-Motivated Approach for Detecting Strategic Adaptation

In many settings, machine learning models may be used to inform decisions that impact individuals or entities who interact with the model. Such entities, or agents, may game model decisions by manipulating their inputs to the model to obtain better outcomes and maximize some utility. We consider a multi-agent setting where the goal is to identify the"worst offenders:"agents that are gaming most aggressively. However, identifying such agents is difficult without knowledge of their utility function. Thus, we introduce a framework in which each agent's tendency to game is parameterized via a scalar. We show that this gaming parameter is only partially identifiable. By recasting the problem as a causal effect estimation problem where different agents represent different"treatments,"we prove that a ranking of all agents by their gaming parameters is identifiable. We present empirical results in a synthetic data study validating the usage of causal effect estimation for gaming detection and show in a case study of diagnosis coding behavior in the U.S. that our approach highlights features associated with gaming.

Paper

References (63)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer Drhf6/10 · confidence 2/52024-07-11

Summary

The paper studied how to identify and rank agents who strategically manipulate their inputs to game machine learning models in multi-agent settings. A causally-motivated approach was proposed to address this challenge.

Strengths

The paper is an interesting follow-up to [1]. Provided that the proposed mythology is solid, I can foresee many scenarios where it can be applied. Especially, I feel that it may be applied to analyze college admission mechanisms. [1] Hardt, Moritz, et al. "Strategic classification." Proceedings of the 2016 ACM conference on innovations in theoretical computer science. 2016.

Weaknesses

The paper assumes players to be almost perfectly rational, i.e., they have a strong motivation to maximize utility. In reality, however, people are often only bounded rational. If so, that is, if some players do not particularly care about their utility, how would the accuracy of your approach be affected?

Questions

Please find my first question in the above part. Here's another question. After the proposed approach is adopted and the player gaming the most is identified, how can we find our if the identification is correct? That is to say, how to validate your approach.

Rating

6

Confidence

2

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors have appropriately discussed the limitations.

Reviewer P2tW6/10 · confidence 2/52024-07-12

Summary

This work studies the problem of identifying agents who would likely game a given system. When the gaming parameters are unknown, the authors show that identifying these parameters requires strong assumptions. In contrast, they show an ordering of agents based on their ranking order is learnable from a dataset. They use this ranking to detect gaming in a Medicare application.

Strengths

The authors generalize the study of strategic adaptation to a much more realistic setting and provide provable results for computing rankings, which can then be used to detect gaming. The application of Medicare is also interesting.

Weaknesses

Although the framework is interesting, it is unclear in general how to use rankings to detect gaming, which is the motivator of the study. A general provable approach to using rankings for gaming, even under some conditions, would improve the applicability of this work.

Questions

Is there any general approach to using rankings as a subroutine to detect gaming? If so, under what conditions would this approach provably work?

Rating

6

Confidence

2

Soundness

4

Presentation

2

Contribution

3

Limitations

Yes.

Reviewer dPj83/10 · confidence 3/52024-07-13

Summary

The paper considers the problem of identifying agents with the highest values of scaling parameters in a stylized strategic adaptation optimization model under a wide range of assumptions. The paper casts this problem as a ranking problem via causal effect estimation and provides an algorithm to rank the parameters of agents. The paper then provides an experimental evaluation using synthetic and real-world medicare data to show the observations from the proposed models and algorithm.

Strengths

+ The paper considers an interesting manipulation problem

Weaknesses

- The motivation and justification for the considered problems seem to be lacking at the beginning (mostly presented with very little context and connections) - The model is very marginal compared to existing studies (i.e., the new component seems to assume non-similarities in gaming via a cost-scaling factor); the notations/paper contents are poorly written, the results have few explanations in terms of implications, and there are assumptions in the models that are hard to justify - The proposed approaches (i.e., casts as causal effect estimation) have little to no justifications - The synthetic experimental evaluation is based on a single dataset with manual tuning that has little to no explanations, which limits the generalizability of the observations

Questions

N/A

Rating

3

Confidence

3

Soundness

2

Presentation

2

Contribution

2

Limitations

See the above and below comments. Abstract "machine learning models" -> provide examples of how the interaction would look like "their inputs to the model" -> such as? what are the outcomes here? from the ML models? what utility? the agents using the models for something? it is very unclear "We consider a multi-agent setting" -> what is the motivation for this goal? what is the implication if it is achieved? why do we care if it games? What do you mean by aggressively? in fact, how do you even define aggressive? "identifying such agents is difficult" -> why? it is not clear to me how would you do this even if you know the utility; why make it more difficult now? " is parameterized via a scalar" -> the meaning is not clear; what is parametrized? agent utility function? what is a scalar here? why is this realistic at all? "is only partially identifiable" -> what does it mean by this? "By recasting the problem" -> why recasting the problem? why this approach is justifiable? "causal effect estimation problem" -> what is this problem? what is the connection to the worst offenders? very confusing here; the next phrase makes little to no sense at all; why/what is the identifiable property? "in a synthetic data study" -> only one data set? what is this coding behavior? what does it have to do with gaming? what are the features? what do you mean associated with gaming? 1 Introduction "guide decisions that impact individuals or entities" -> provide examples please review the literature on mechanism design and game theory as well as their ML related applications "obtain a more desirable outcome" -> outcome is what here? what is the connection between ML models and outcomes? "to the difficulty of generating supporting evidence" -> provide examples; explain more ""utility maximization:" -> fix "to maximize a payout" -> what is the payout? utility function? "which calculates a" -> the government or the companies? " via a publicly available model" -> how do you know they use this model? "Companies may attempt to" -> what happens to the companies that do this? "Beyond health insurance, gaming emerges" -> how does health insurance impact individuals? who is doing the gaming in the said applications? what are you trying to change in the features? Also, are the models already trained? or are you talking during learning? "agents with the highest propensity " -> how do you measure this? "given a dataset of agents" -> dataset of agents doesn't sound right? what does it mean here? do you mean a set of agents? "their observed model inputs" -> is the model fixed here? how do they interact with the models? why can't do this together? "fraud/gaming labels" -> what does this mean? what is the utility here? what are they affecting? what is the connection to distributional assumptions? these are used with little to no context "But past works in strategic adaptation" -> what does it mean by costs in this context? is that the only difference from this vs the other models? "scalar gaming parameter that scales costs " -> the meaning of this is not clear; what is this parameter? it is not explained in texts; the figure is not meaningful without a clear description "partially identifiable" -> what is the meaning of this? "However, by recasting gaming detection as causal" -> again, why is this justifiable? what is this causal effect estimation going to do? "that ranking agents" -> what ranking is solving? I am very confused "a cost-scaling factor" -> why is this more realistic? "Furthermore, much work in strategic a" -> this paragraph seems to be more like related work section; it is out of place and disrupts the flow of reading; you should probably consider an independent related work section; they are also hard to understand with little to no background context presented What is the contribution here? there should be a contribution subheadings and/or related work also "a synthetic dataset" -> how do you have a ground truth on this? what about other datasets? "causal approaches rank the " -> hard to understand this sentence "healthcare providers, a suspected driver of gaming" -> you can say that unless you have concrete evidence; otherwise, you will sued with claims like this "In summary: we " -> this paragraph provides new information that is not discussed or connected to earlier messages 2 Background & Problem Setup "to a payout" -> what is a payout here? why f is mapping only a single agent attribute? should it maps from more agents? or even datasets? "according to some function" -> what is this function? How is R dependent on d and f? why d' has to be in D? "For simplicity, we assume R = f ; " -> if R = f, the function in (1) makes no sense what is the meaning of f has itself as input? the math is not correct How is strategic classification used in strategic adaptation? "To extend strategic adaptation to multiple agents" -> why is this justifiable at all? What is M_p? What is d_i modeling? what decisions are they making? "Agent assignment is" -> how is this model? how do you indicate assumption; i am still not clear about D_p? is that for an agency? within the agnecy they have M_p agents who can manipulate? " to obtain a higher payout." -> why do they want to increase d_i? why can't they decrease? What do they perturb the average instead of individual d_i? how does this connect to f defined earlier? "is the ground truth value" -> how do you actually get the ground truth? before you define c(,) as two parameters and now you have a single parameter; it doesn't look consistent "we introduce assumptions on the" -> you need to justify these assumptions; are they common in the literature? how do you model multiple agents here? it seems (2) only for one agent 3 Theoretical analysis: finding agents most likely to game "We aim to identify agents most likely" -> wait; is lambda_p unknown? why is this the right way to identify agents? "be point-identified" -> unclear meaning what is the meaning of partially identifiable? You need to provide implications of Prop 1 I still don't get why estimating counterfactuals is used in this context; how come figure 2 has no in-text explanations? There are algorithms and figure on page 5 without explanations and connections to the paper 4 Empirical results & discussion "We hand-select " -> what about other connection

Authorsrebuttal2024-08-06

Rebuttal by Authors (part 2/2)

**Connections to past work.** The reviewer suggests connections to the mechanism design literature. We disagree: our proposed framework differs from mechanism design. Algorithmic mechanism design aims to maximize some social surplus (e.g., sum of utilities for all agents, in the language of our framework) by learning an optimal allocation rule (e.g., what we call a “model”) of some “resource” to agents [A3]. In contrast, our model/allocation rule is fixed, and we study how multiple individual agents respond. This aligns with our motivating problem setting in U.S. Medicare: given a fixed model (CMS-HCC) for calculating payouts to healthcare providers/insurance companies, providers/companies adjust their behavior accordingly. A mechanism design approach could inform upgrades to the CMS-HCC model, but is not our focus. We believe that our review of game-theoretic work in machine learning is extensive, though we are happy to consider suggestions to discuss specific papers. To that end, the closest related area is strategic classification, in which agents maximize their utility function in response to an ML model’s decisions. To that end, we have provided citations to related settings in L34-L38. We pay special attention to works with similar assumptions to ours, *e.g.*: * Utility functions are partially known (their setting) [A4] vs. utility functions are unknown beyond the general form (our setting) * Multiple agents with different capacities to game, modeled by norm constraints on manipulation (their setting) [A5] vs. different cost-scaling factors (our setting) We will consider a separate section/heading for the related works. **Model is “very marginal.”** We disagree that this is a weakness: small changes to the problem formulation may significantly affect the solution space. However, in addition to a novel utility-maximization formulation for gaming in multiple agents, we contribute theoretical analyses of the resultant utility-maximization problem, which motivates a causal effect estimation approach to producing a ranking of agents by gaming propensity. To our knowledge, our approach is the first to adopt a causal effect estimation approach to gaming ranking (though previous works have explored causal mechanism design approaches to counter gaming [A6, A7]). **Evaluation on one synthetic dataset => limited generalizability?** We partially disagree: a synthetic dataset for evaluation is necessary to establish proof-of-concept of the proposed approach. This is because ground truth gaming labels are inherently difficult/infeasible to obtain in practice (e.g., require accurate audits of all agents, or accurate self-reporting of fraud). Synthetic data validation is standard in causal inference (e.g., [A8, A9] are two well-established causal effect estimators that use IHDP [A10], a synthetic dataset, for validation). To further mitigate generalizability issues, we verify our rankings align with known drivers of upcoding (e.g., the state-level prevalence of private healthcare providers [A11, A12], Section 4.3, Table 1, pg. 9). We also align our design choices with our problem assumptions (Appendix C.1, pg. 16-17). We are happy to consider specific suggestions for other datasets or design choices that could improve the generalizability of our findings. **Are assumptions common in the literature? How are they justified?** We justify our assumptions after their introduction on L92-110, with examples where applicable. We like the suggestion to discuss how common our assumptions are in the literature; in particular, many strategic classification works implicitly use some form of Assumption 1 [A6, A13], and our Assumptions 2-4 are more general cases of assumptions found in [A4, A13, A14]. Assumption 5 ensures that the ground truth is predictable from the observed variables $x_i$, and is related to assumptions of no unmeasured confounding [A15]. Assumptions 6-8 are standard assumptions in the causal inference literature (e.g., Chapter 3 of [A15]). These hold independently for all agents. We will add these points to the revision. **Other changes.** Thank you for your diligence in checking our notation — we appreciate the careful notes! We will double-check that all variables are well-defined and identify places where we’ve used similar notation/overloaded notation.

Authorsrebuttal2024-08-06

References cited in rebuttal to Reviewer dPj8

**References**\ [A1] Report to Congress: Risk Adjustment in Medicare Advantage, December 2021. https://www.cms.gov/files/document/report-congress-risk-adjustment-medicare-advantage-december-2021.pdf\ [A2] Ding, Peng. A first course in causal inference. CRC Press, 2024.\ [A3] Roughgarden, Tim. Twenty lectures on algorithmic game theory. Cambridge University Press, 2016.\ [A4] Dong, Jinshuo, et al. "Strategic classification from revealed preferences." Proceedings of the 2018 ACM Conference on Economics and Computation. 2018.\ [A5] Shao, Han, Avrim Blum, and Omar Montasser. "Strategic classification under unknown personalized manipulation." Advances in Neural Information Processing Systems 36 (2024).\ [A6] Bechavod, Yahav, et al. "Gaming helps! learning from strategic interactions in natural dynamics." International Conference on Artificial Intelligence and Statistics. PMLR, 2021.\ [A7] Horowitz, Guy, and Nir Rosenfeld. "Causal strategic classification: A tale of two shifts." International Conference on Machine Learning. PMLR, 2023.\ [A8] Shi, Claudia, David Blei, and Victor Veitch. "Adapting neural networks for the estimation of treatment effects." Advances in neural information processing systems 32 (2019).\ [A9] Louizos, Christos, et al. "Causal effect inference with deep latent-variable models." Advances in neural information processing systems 30 (2017).\ [A10] Hill, Jennifer L. "Bayesian nonparametric modeling for causal inference." Journal of Computational and Graphical Statistics 20.1 (2011): 217-240.\ [A11] Silverman, Elaine, and Jonathan Skinner. "Medicare upcoding and hospital ownership." Journal of health economics 23.2 (2004): 369-389.\ [A12] Silverman, Elaine, and Jonathan S. Skinner. "Are for-profit hospitals really different? Medicare upcoding and market structure." (2001).\ [A13] Hardt, Moritz, et al. "Strategic classification." Proceedings of the 2016 ACM conference on innovations in theoretical computer science. 2016.\ [A14] Levanon, Sagi, and Nir Rosenfeld. "Strategic classification made practical." International Conference on Machine Learning. PMLR, 2021.\ [A15] Robins, James, and Hernan, Miguel A. “Causal Inference: What If?” Boca Raton: Chapman & Hall/CRC. (2020).

Reviewer P2tW2024-08-12

Rebuttal Response

Thank you for your response! Although the ranking approach is interesting, I will stick with my given scores since the main goal of the paper is to detect gaming. If your preliminary idea works out, it would greatly strengthen the paper.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC