Thank you very much for your reply and the detailed explanations. However, I remain unconvinced about the novelty of "connecting all these ideas to the ELBO", as well as the claim that the BONG update's "derivation and application to variational inference are novel".
While it is true that Lyu & Tsang's paper "goal is optimization, not inference or Bayesian updating", it was merely an example of a paper that provides a general analysis of this specific algorithm. The idea has been explored in various papers with different approaches and motivations, such as [van der Hoeven et al., 2018; Cherief-Abdellatif et al., 2019; Khan & Rue, 2023]. In particular, the authors of [Khan & Rue, 2023] discuss the online VI case at the end of Section 5 and mention the linearization of the expected loss in the ELBO, referring to [van der Hoeven et al., 2018], which delves into these questions in detail (though not through the lens of VI). From a different perspective, [Cherief-Abdellatif et al., 2019] establish a connection between VI and online learning by proposing to perform a single gradient descent step for the expected loss without the KL, which they show to be equivalent to minimizing the ELBO with a linearized expected loss.
Therefore, while I do believe this is a nice piece of work, I am still not fully convinced of its novelty. And although I recognize that the paper offers more than just this point, I feel that it is the aspect most emphasized (particularly in the abstract).
[van der Hoeven et al., 2018]: D. van der Hoeven, T. van Erven, and W. Kotłowski. The many faces of exponential weights in online learning. In Conference On Learning Theory, pages 2067–2092. PMLR, 2018.
[Cherief-Abdellatif et al., 2019]: B.E. Cherief-Abdellatif, P. Alquier, and M.E. Khan. A generalization bound for online variational inference. In Asian Conference on Machine Learning, pages 662–677. PMLR, 2019.
[Khan & Rue, 2023]: M.E. Khan and H. Rue. The Bayesian Learning Rule. In Journal of Machine Learning Research, 2023.