Summary
The paper builds on the duality between estimation and control, in order to develop an online data-driven method for MSE-optimal filtering of linear systems with linear observations. Process and observation noise covariances are considered unknown, but stochastic states are assumed to be bounded. The paper proposes using SGD in the space of steady-state stabilizing gains, claims asymptotic convergence to the optimal gain, and provides an asymptotic probabilistic bound on the deviation from optimal error.
Strengths
Although the use of online optimization for controlling linear-quadratic settings is not new (e.g. [r1]), exploiting the duality between control and filtering problems is novel and interesting in this context. The concentration and error bounds are not trivial and useful, as concentration bound is non-asymptotic in the series length T.
Proof of SGD convergence and error bounds are novel as well, even if derived under very strong assumptions.
[r1] Cohen A. et al., "Online Linear Quadratic Control", 2018
Weaknesses
Overall, I believe the authors did their best that the paper will be well-organized and clear (as it can be for a rather technical manuscript). The explanations and given outlines before each section are indeed helpful. However, it is still very hard to follow the assumptions and constant definitions, and some of the conclusions. Writing becomes very laconic at some crucial points (e.g. Thm. 2, Remark 7).
Particularly, Thm.2 and it's proof are not clear hence it is hard to get convinced in their soundness. Furthermore, In the introduction it is stated that convergence is guaranteed from every initial policy, but according to Theorem 2 the policy cannot enter (or start in) a class of policies where gradient is smaller than some constant. I didn't find any discussion about when trajectories enter this region. My impression is that this is a good paper, and I might be missing something. I will be willing to raise my score when given a more detailed proof and this clarification.
Questions
I ask for a more detailed proof for Thm. 2 (see above).
In addition, a summary of all assumptions, results and notations can be very useful.
Minor typos:
l. 249 't' should be replaced by $\gamma$.
l. 276 I think i should go between 1 and T-1.
l. 309 quite.
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
The online method uses a surrogate loss using the observations, this is a reasonable choice due to lack of ground-truth states and the observability, which is a strong assumption. The paper should discuss the implications of using this loss with more general (i.e. non observable) systems.