Distributionally Robust Linear Quadratic Control

Linear-Quadratic-Gaussian (LQG) control is a fundamental control paradigm that is studied in various fields such as engineering, computer science, economics, and neuroscience. It involves controlling a system with linear dynamics and imperfect observations, subject to additive noise, with the goal of minimizing a quadratic cost function for the state and control variables. In this work, we consider a generalization of the discrete-time, finite-horizon LQG problem, where the noise distributions are unknown and belong to Wasserstein ambiguity sets centered at nominal (Gaussian) distributions. The objective is to minimize a worst-case cost across all distributions in the ambiguity set, including non-Gaussian distributions. Despite the added complexity, we prove that a control policy that is linear in the observations is optimal for this problem, as in the classic LQG problem. We propose a numerical solution method that efficiently characterizes this optimal control policy. Our method uses the Frank-Wolfe algorithm to identify the least-favorable distributions within the Wasserstein ambiguity sets and computes the controller's optimal policy using Kalman filter estimation under these distributions.

Paper

Similar papers

Peer review

Reviewer iQxV7/10 · confidence 3/52023-07-01

Summary

The paper **Distributionally Robust Linear Quadratic Control** considers the case of finite-time linear quadratic optimal control with process and measurement noises and known time-varying dynamics. The novelty lies in the fact that the distributions of the initial state and noises are unknown but lie in an ambiguity set defined as a 2-Wasserstein ball centered in a known Gaussian distribution and with known radius. The objective is to minimize the expected LQR cost with the distributions chosen adversarially in that ball. The authors call this adversarial problem "distributionally robust Linear Quadratic Gaussian". It strict generalizes LQG control, which assumes these distributions are known and Gaussian. The main contribution is twofold. The first part is theoretical, as the authors prove the existence of a Gaussian adversarial distribution, that is, the distributionally robust LQG reduces to classic LQG with unknown Gaussian distributions. A consequence is that the optimal controller is a linear state feedback, like in the classic LQG case. This leads to the second contribution: an algorithm to efficiently estimate the adversarial Gaussian distribution from data. The authors claim that, then, the LQG problem with estimated noise distributions can be solved efficiently by using classical methods. The authors illustrate convergence of their estimation algorithm on a simulated example.

Strengths

The paper is interesting and the clear exposition makes the reasonings easy to follow. The appendix is well-managed, and I appreciate Appendix A on the solution of classical LQG with a Kalman filter. I like the idea of allowing for a whole family of noise distributions rather than assuming a fixed one. The main theoretical result helps mitigate the restrictiveness of the commonly-accepted "white noise assumption"; indeed, it shows that allowing for a larger family of noises does not provide any benefits (for distributions close enough to Gaussians). This new problem formulation thus seems relevant. The subsequent algorithm to solve the distributionally robust LQG problem follows naturally and further justifies the interest of the theoretical result. Overall, the reasoning exposed is sound and well-motivated. Finally, I appreciate that the authors went the extra mile in Section IV by augmenting the theoretical result with a data-efficient algorithm.

Weaknesses

The three main weaknesses of the paper are, in my opinion, 1. the lack of thoroughness of the simulation study; 2. the lack of discussion of the effect of hyperparameters; and 3. some arguments are unclear and should be explicited (although I do not question their conclusions). I detail these three points. These points are not critical for acceptance, but I believe the paper would be improved by addressing them. W1) The simulation study shows convergence of the algorithm on a randomized use case. This is a good sanity check, but I would also appreciate a simulation of the case when the noise distribution is not normal but still within the Wasserstein ball. The theoretical result ensures that Nature's adversarial strategy _is_ normal, but an illustration of the case when Nature is sub-adversarial would be welcome. In particular, is the cost of the learned policy reduced compared to when the noise is normal? W2) The role of the choice of ball radii is not discussed. While there is obviously no notion of "optimal" radius, since it simply corresponds to a degree of robustness, I would like to see the evolution in performance with increasing radii. Intuition tells that performance should decrease as robustness increases, but a confirmation in simulation would be welcome. W3) Some arguments are unclear to me, although their conclusions seem to be valid. In particular: 1. The authors make multiple claims on convexity of sets of probability measures. Examples are lines 112, 151, 189. I understand that these claims are made by considering Borel measures as a subset of the vector space of signed Borel measures with standard addition and scalar multiplication. I believe this superset should be mentioned explicitly at least ones, as it is rather unusual and notions of convexity of a metric space exist (and differ). This is also relevant on line 151, where $\mathcal{W}$ is claimed to be infinite-dimensional despite not being a vector space. 2. With this understanding of convexity, why is the set $\mathcal{W}$ non convex, as claimed e.g. on lines 112, 151 and 189? As far as I understand, each set $\mathcal{W}_{z}$ is convex, with $z\in\{x_0, w_t, v_t\}$. Then, $\mathcal{W}$ should be convex as the cartesian product of these sets. The only way I understand this non-convexity is if $\mathcal{W}$ is not _equal_ to the cartesian product, but only isometric to this cartesian product by the mapping that computes the marginal distributions. I believe this should be mentioned explicitly, at least in a footnote, as this questions distracted me from more central claims of the paper for a while. 3. I am unsure about the claim on line 131 that the controller can compute the fictious states $\hat x_0,\dots,\hat x_t$ from the real observations $y_0, \dots, y_t$ without knowing the initial state. Indeed, this would imply in particular that the controller can reconstruct the initial state. Since the time origin is arbitrary, any state could be reconstructed. This claim appear unnecessary for the rest of the argument and should be either removed or clarified in my opinion. 4. I find the formulation of the first sentence of Proposition 4.1 extremely confusing. In particular, what comes after "then" in the first sentence reads as a logical consequence of what precedes whereas it is actually a definition of the symbols $\mathbb{P}^\star$ and $V_t^\star$. I recommend reformulating. 5. In Proposition 4.2, I recommend avoiding using the term "smooth". While it has a precise meaning in a specific branch of mathematics, it is often only used informally in control. I would prefer the more standard "infinitely differentiable with $\beta$-Lipschitz gradient".

Questions

My questions are detailed in the above paragraph on weaknesses. I would appreciate if the authors could respond to these.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors have adequately addressed limitations and potential negative societal impact of their work.

Reviewer BGQp7/10 · confidence 4/52023-07-02

Summary

The paper proposes a robust method for controlling LQ systems, where process and observation noise distribution laws are unknown, but noise samples are assumed to be independent, zero mean and have their distributions lying close to a a nominal Gaussian distribution in Wasserstein-2 space. A pair of relaxed settings is used to prove a strong duality between the minimax\maximin problems, showing that the worst-case distribution is Gaussian. A numerical method is proposed for computing this distribution for long horizons.

Strengths

The paper itself is well-written and gives a thorough overview of the problem. The fact that the solution for the minimax problem is given by a Gaussian distribution \ linear controller, even if not surprising considering the nature of the quadratic cost and $W_2$/Gelbrich distance, is still not clear from the outset and is important enough from a theoretical point of view. Its proof seems correct as well.

Weaknesses

My main concern here, is that the theoretical contribution is limited to one main result, which is confined to a Gaussian-centered ambiguity set, i.e. the real distribution is assumed to be close to a Gaussian one. This calls for an additional study of other nominal distributions (as the Gelbrich distance equals to $W_2$ for other families of distributions [18], and discussion may be confined to linear filters) or mixtures, or at least for a detailed discussion about the practical implications of this choice (beyond that of the 'concluding remarks' in Sec.6) i.e. about whether this hypothesis may (or might not) be practical in real problems. However, even if limited, the contribution is still novel and important enough to warrant acceptance.

Questions

Some minor comments \ typos: It is somewhat misleading to call $\mathcal{W}$ a 'Wasserstein ball', e.g. l.33, while it is not even convex, but I guess that's forgivable since this set is defined clearly. l. 93 dimension should be $p \times T$. l.104 it would probably be more didactic to introduce the Wasserstein distance before using it to define $\mathcal{W}$.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors addressed the limitations, but a more detailed discussion about the assumption that the nominal distribution is Gaussian is required.

Reviewer P3Ek7/10 · confidence 3/52023-07-05

Summary

The paper proposes a distributionally robust version of the output feedback linear quadratic control problem. The goal is to control a system with known partially observed linear dynamics in the face of stochastic disturbances. The stochastic disturbances (both measurement and state disturbances) are drawn from an unknown distribution. This distribution is assumed to belong to some ambiguity set centered at a known gaussian distribution. The disturbances drawn from any distribution in the ambiguity set are assumed to be mean zero, and independent across time, with measurement disturbances independent from the process disturbances. Furthermore, the marginal distribution for the measurement or process disturbance at each time is assumed to be close to the corresponding nominal marginal distribution (as measured by the 2-Wasserstein distance). Under this setting, the paper finds the optimal output feedback controller for the worst case distribution of disturbances in the ambiguity set. It is found that the worst case distribution of disturbances belonging to the ambiguity set described above is Gaussian. Therefore, the corresponding optimal controller is a Linear-Quadratic Gaussian controller designed for this worst case distribution. It is then shown that the worst case distribution may be determined in a computationally efficient manner. In particular, it is demonstrated that the Frank-Wolfe algorithm converges to the optimal covariance parameters for the worst case distribution. Furthermore, it is shown that each step of the Frank-Wolfe algorithm may be decomposed into simple, easily parallelizable components.

Strengths

The formulation of the distributionally robust output feedback LQ control problem is novel. It is related to previously published results on distributionally robust state feedback LQ control. As acknowledged by the authors, the output feedback setting brings an extra technical challenge due to the dependence of the optimal state estimator upon the disturbance distribution. The problem and corresponding solution are clearly presented, and easy to follow. The authors highlight the rather surprising result that the worst case disturbance distribution from the prescribed ambiguity set is Gaussian. Given the appropriate context, the results here could be a meaningful step in a unification of worst case and stochastic control.

Weaknesses

The contextualization of the setting studied relative to prior work could be improved. In particular, there are a large class of control synthesis approaches for settings where the noise distributions are either not available or are not Gaussian. In particular, consider H-infinity approaches [Zhou et. al, Robust and Optimal Control, 1996], mixed H-2/H-infinity approaches [(Doyle et. al Optimal Control with Mixed H-2/H-infinity Performance Objectives, 1989) (Bernstein and Haddad, LQG control with an H-infinity performance bound, 1988)], adversarially robust control [Lee et. al, Performance-Robustness Tradeoffs in Adversarially Robust Control and Estimation, 2023], and nonstochastic control [Hazan and Singh, Introduction to Online Nonstochastic Control, 2023]. It would be useful to discuss several of these results and contrast the setting with the distributionally robust setting. A thorough discussion about the choice for the ambiguity set and what distributions it can model would be beneficial. In particular, considering mean zero disturbances which are independent across time is restrictive. It fails to model e.g. colored noise. Despite the fact that it was possible to propose a relatively efficient method to optimize for the worst-case covariance of the noise distribution, the required computation time appears to scale quite poorly with the problem horizon, T. Only short horizons, up to T=20, were considered in the simulations, however many control problems have much longer horizons. This appears to limit the practical applicability.

Questions

Is there any concrete theoretical connection between distributionally robust linear quadratic control and the other methods for incorporating robustness mentioned above (e.g. mixed H-2/H-inf)? Are there any practical examples where the proposed method substantially outperforms conventional approaches for incorporating robustness to unknown disturbance distributions? Such an example would make the setting more compelling for practical use. From the experiments, is it possible to detect any clear trends regarding the worst case covariance relative to the nominal covariance of the ambiguity set? E.g. Do we see or expect to see that the worst-case covariances are larger than the central covariances in Loewner order? Combined with the theoretical results, such an observation might justify crude approximate approaches for distributionally robust control design in practice. For example, one could take their estimate for the covariance for the central distribution in the ambiguity set, and simply scale it up by some constant. The resulting covariances could then be used to design a LQG controller.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

2 fair

Limitations

There is no negative societal impact that the authors must address. Limitations to the method that would be worth addressing are mentioned in the weaknesses section. To reiterate, it would be helpful to address: - What the ambiguity set can model, and what it cannot - In which settings the proposed method would outperform conventional methods from robust control theory for handling unknown disturbance distributions - Acknowledging the computational burden of the approach relative to e.g. LQG with a known distribution

Reviewer RqGY7/10 · confidence 4/52023-07-08

Summary

- The paper considers a standard LQ setup with uncertainty in the distributions of system noise, observation noise and initial state. - Described as Zero-Sum Game between the controller and nature, they show the optimal decisions for both players. Specifically they show that the worst-case distribution is Gaussian and the optimal control law is linear. - They provide an numerically efficient approach to solve the distributionally robust LQ control problem based on Frank-Wolfe algorithm. - Simulation experiments are provided to show the computational efficiency of the proposed algorithm compared to MOSEK.

Strengths

The results provided in the manuscript make multiple important contribution. - The optimal control law still remains linear and the worst case distribution is still Gaussian. - The proof technique is novel which relies on the "purified states" instead of usual dynamic programming approaches, it is interesting to see how this approach can be used in other LQG problems. - Simulation results show the computational efficiency of the Frank-Wolfe algorithm over MOSEK.

Weaknesses

- There are no obvious major weakness in the manuscript. Some potential minor weakness are currently mentioned in the form of questions in the comment below.

Questions

- In general the adaptive linear quadratic control results are provided under assumption of sub-gaussian system noises (Assumption A1, http://proceedings.mlr.press/v19/abbasi-yadkori11a/abbasi-yadkori11a.pdf). Is it possible to extend these results to Sub-Gaussian nominal distribution?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Authors provide some limitations in form of possible extensions and future work in the last section. Social Impact: NA

Reviewer RqGY2023-08-10

Response to the Rebuttal

I thank the authors for their responses. After reading their rebuttal, I will retain my recommendation regarding this paper.

Reviewer BGQp2023-08-13

I thank the authors for their detailed response. I am excited to see you managed to improve your results. Although I believe the new results are sound and extend the theory, it is hard to judge since they haven't been reviewed. I will retain my recommendations.

Reviewer P3Ek2023-08-16

Thank you - your response mostly addresses my questions. While I will raise my score, I do want to see better physical justification for the chosen ambiguity set.

Authorsrebuttal2023-08-17

Thank you very much for raising your score. We will provide additional justification for using Wasserstein ambiguity sets in the final version of the paper.

Reviewer iQxV2023-08-16

Thank you for the response and clarifications. I have no further questions at this stage and will retain my score.

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC