Summary
This paper investigates the effectiveness of decomposition approaches for solving risk-averse Markov Decision Processes (MDPs) with Conditional Value-at-Risk (CVaR), Expected Value-at-Risk (EVaR), and Value-at-Risk (VaR) objectives. The goal is to assess the validity and accuracy of these common decomposition techniques for risk-averse decision-making.
Main contributions:
Sub-optimality of CVaR and EVaR decomposition: The paper demonstrates that the widely used decomposition approaches for CVaR and EVaR objectives are inherently suboptimal, invalidating previous works. These methods involve a saddle-point gap during policy optimization, leading to incorrect results. As a result, practitioners should exercise caution when using these decomposition techniques for risk-averse MDPs.
Optimal VaR decomposition: In contrast to the suboptimal CVaR and EVaR decompositions, the paper identifies the VaR decomposition as an optimal approach for both policy evaluation and optimization in risk-averse MDPs. The VaR decomposition does not suffer from the same saddle-point problem, making it a reliable method for risk-averse decision-making.
Increased awareness and call for alternative approaches: The findings highlight the limitations of two traditional decomposition methods, CVaR and EVaR, and emphasize the need for scrutiny when using these decompositions. Researchers are encouraged to explore alternative approaches, such as parametric dynamic programs, to improve risk-averse decision-making in MDPs.
Strengths
Rigorous analysis: The main strength of the paper is its rigorous treatment of three decomposition approaches for risk-averse MDPs.
Significant and clearly stated contribution for the field: The paper shows the sub-optimality of CVaR and EVaR decomposition methods invalidating several previous works on the topic. This finding is valuable as it warns practitioners about unknown limitations of these techniques in risk-averse decision-making.
Optimal VaR decomposition: By identifying the VaR decomposition as an optimal approach for policy evaluation and optimization in risk-averse MDPs, the paper provides a practical solution to the limitations of CVaR and EVaR. Thus this VaR decomposition is a reliable alternative to these two methods.
Encouraging further research: In light of the limitations faced by CVaR and EVaR, the paper calls for exploring alternative approaches, such as parametric dynamic programs.
Clarity in presentation and methodology: The paper is well-written. It uses clear explanations, comprehensive mathematical analysis, and compelling illustrative examples.
Weaknesses
Limited set of approaches: Perhaps the main weakness of the paper is the limited set of (three) decomposition approaches being scrutinized. While the paper highlights the sub-optimality of CVaR and Expected Value-at-Risk EVaR methods, it does not compare a broader range of state-of-the-art techniques. This limitation makes unclear the extent to which VaR decomp Google Mapsosition can be considered optimal. I appreciate that it would not be possible, given length constraints and scope of the paper, to evaluate other baselines comprehensively. However, it would be valuable if the authors could mention other state-of-the-art techniques in a related works section, and whether they can say anything about them in relation to the approaches they investigated here. For instance, control as inference and active inference use the KL divergence as a decision-making objective (i.e., a belief-based reward function) which entails risk-averse behaviour, and which might be related to VaR [1,2,3].
Unclear scalability discussion: The paper does not extensively discuss the scalability of the proposed VaR decomposition concerning problem size or dimensionality. I appreciate that this might not be important for theoretically minded readers, however, since the paper has important implications for practitioners (i.e., widespread adoption of VaR) it would be nice to immediately know, via some discussion or suitable reference, how well this method scales as MDPs grow complex and large in applications.
Limited generalization: The paper comprehensively addresses the setting of risk-averse MDPs. It would be nice to know whether the theoretical analyses presented here can offer some insights on more general problem classes like risk averse partially observable Markov decision processes (POMDPs). For example, has VaR been extended to risk averse POMDPs, and if so would you expected it to be optimal in this case?
From the two points above, one of the main weaknesses of this paper is that it is not straightforward to infer to what extent VaR is a viable approach in a wide range of practical problems, and thus it is unclear to what extent the paper implies an optimal, promising in the near long-term approach to the risk-averse decision problem. In my understanding, this is the main point of broader impact in the machine learning community, so it should be addressed.
Typos: l 136, l200
[1] S. Levine, ‘Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review’, arXiv:1805.00909 [cs, stat], May 2018, Accessed: Dec. 29, 2021. [Online]. Available: http://arxiv.org/abs/1805.00909
[2] L. Da Costa, N. Sajid, T. Parr, K. Friston, and R. Smith, ‘Reward Maximization Through Discrete Active Inference’, Neural Computation, vol. 35, no. 5, pp. 807–852, Apr. 2023, doi: 10.1162/neco_a_01574.
[3] D. Hafner, P. A. Ortega, J. Ba, T. Parr, K. Friston, and N. Heess, ‘Action and Perception as Divergence Minimization’, arXiv:2009.01791 [cs, math, stat], Oct. 2020, Accessed: Nov. 07, 2020. [Online]. Available: http://arxiv.org/abs/2009.01791
Questions
Can the authors elaborate on their selection of the specific decomposition approaches used for comparison? Are there any reasons why these three approaches were chosen over others?
Could the authors provide more insights into the computational complexity and scalability of the VaR decomposition method concerning the size and dimensionality of the risk-averse MDPs, either in a short discussion or references? Are there any notable challenges when applying this approach to large and more complex decision problems?
As a broader point of interest, how significant is the choice of the risk level (alpha) in the performance and efficiency of the VaR decomposition method? Are there any guidelines on selecting appropriate alpha values based on a given problem characteristics? This is important to understand the robustness of this method in a new risky environment, a desideratum for any risk averse decision-making algorithm in high stakes applications, which relates to the broader impact of this work as hinted in the last point of the conclusion.
Given that the paper focuses on risk-averse MDPs, are there any indications or early results on how the VaR decomposition approach and its optimality might generalize to other types of decision-making frameworks, such as partially observed problems?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
The paper does a good job at addressing its limitations insofar as it focuses on three specific approaches to risk averse decision making.
The main limitation which has not been addressed, and which in my opinion should be addressed, is giving more context as to why this three approaches were chosen, or why is it is sensible to choose them, and mention the fact that there exist other state-of-the-art approaches, which need to be considered and comprehensively evaluated in the future, and which could serve as avenues for future research in addition to parametric dynamic programs, e.g., control as inference and active inference.