Summary
Contextual stochastic bilevel optimization (CSBO) is introduced in this paper. An efficient double-loop gradient method based on the Multilevel Monte-Carlo (MLMC) is proposed. The proposed framework captures important applications such as meta-learning, Wasserstein distributionally robust optimization with side information (WDRO-SI), and instrumental variable regression (IV).
Strengths
An interesting problem, that is, Contextual Stochastic Bilevel Optimization Problem is proposed in this work. And an efficient algorithm is proposed to solve this problem. The proposed framework captures important applications such as meta-learning, Wasserstein distributionally robust optimization with side information (WDRO-SI), and instrumental variable regression (IV). However, I have some concerns as follows.
Weaknesses
1. I think it's necessary to emphasize the difficulty of solving the Contextual Stochastic Bilevel Optimization Problem compared with traditional bilevel optimization problems. The presentation of this work is poor. I believe this is an excellent work, I suggest that you should modify the presentation of this work to better clarify the contributions of this work.
2. In the experiment, the results are limited. For example, in the meta-learning application, I suggest the authors compare the proposed method with the state-of-the-art bilevel optimization methods [1][2][3], which are shown to be able to address the meta-learning task. Alternatively, you need to specify why these methods [1][2][3] are not applicable to this application. Furthermore, the authors should conduct experiments on more datasets to better evaluate the proposed method, for example, Omniglot dataset.
3. I suggest the author briefly introduce some existing bilevel optimization works in machine learning, for example, hyper-gradient-based methods [1, 4] and approximation-based methods [2], and then discuss why these methods fail to be applied to the Contextual Stochastic Bilevel Optimization Problem.
[1] Bilevel optimization: Convergence analysis and enhanced design, ICML, 2021
[2] Asynchronous Distributed Bilevel Optimization, ICLR 2023
[3] Bilevel Programming for Hyperparameter Optimization and Meta-Learning, ICML 2018
[4] Provably faster algorithms for bilevel optimization, NeurIPS 2021
Questions
1. I think it's necessary to emphasize the difficulty of solving the Contextual Stochastic Bilevel Optimization Problem compared with traditional bilevel optimization problems. The presentation of this work is poor. I believe this is an excellent work, I suggest that you should modify the presentation of this work to better clarify the contributions of this work.
2. In the experiment, the results are limited. For example, in the meta-learning application, I suggest the authors compare the proposed method with the state-of-the-art bilevel optimization methods [1][2][3], which are shown to be able to address the meta-learning task. Alternatively, you need to specify why these methods [1][2][3] are not applicable to this application. Furthermore, the authors should conduct experiments on more datasets to better evaluate the proposed method, for example, Omniglot dataset.
3. I suggest the author briefly introduce some existing bilevel optimization works in machine learning, for example, hyper-gradient-based methods [1, 4] and approximation-based methods [2], and then discuss why these methods fail to be applied to the Contextual Stochastic Bilevel Optimization Problem.
[1] Bilevel optimization: Convergence analysis and enhanced design, ICML, 2021
[2] Asynchronous Distributed Bilevel Optimization, ICLR 2023
[3] Bilevel Programming for Hyperparameter Optimization and Meta-Learning, ICML 2018
[4] Provably faster algorithms for bilevel optimization, NeurIPS 2021
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.