Bayesian Adaptive Calibration and Optimal Design

The process of calibrating computer models of natural phenomena is essential for applications in the physical sciences, where plenty of domain knowledge can be embedded into simulations and then calibrated against real observations. Current machine learning approaches, however, mostly rely on rerunning simulations over a fixed set of designs available in the observed data, potentially neglecting informative correlations across the design space and requiring a large amount of simulations. Instead, we consider the calibration process from the perspective of Bayesian adaptive experimental design and propose a data-efficient algorithm to run maximally informative simulations within a batch-sequential process. At each round, the algorithm jointly estimates the parameters of the posterior distribution and optimal designs by maximising a variational lower bound of the expected information gain. The simulator is modelled as a sample from a Gaussian process, which allows us to correlate simulations and observed data with the unknown calibration parameters. We show the benefits of our method when compared to related approaches across synthetic and real-data problems.

Paper

Similar papers

Peer review

Reviewer CjgP7/10 · confidence 3/52024-07-10

Summary

This paper addresses the problem of calibrating simulation models. Simulation models depend on inputs set by the user, referred to as designs, and parameters representing unknown physical quantities, called calibration parameters. The task is to find calibration parameters such that simulations match real observations. To that end, the authors propose an active learning scheme in which maximally informative design and calibration parameters are iteratively chosen to construct the training set. They also propose a Gaussian process structure adapted to this setting.

Strengths

### Originality * This piece of work is new to me. ### Quality * The method is sound, and I did not identify any flaws. * Related works are discussed. * The experiments support the claims. ### Clarity * The paper is well articulated. * The problem is clearly introduced, and the intuition behind the proposed solution is provided early on. * Figures are clean. ### Significance * Although my knowledge of the field is too limited to have a strong opinion, the problem addressed seems important.

Weaknesses

### Originality * I have no concerns regarding originality. ### Quality * I have no concerns regarding the quality. ### Clarity * Calibration parameters are also referred to as simulation parameters, which is confusing. * Some minor comments for the camera-ready version. There is a typo in equation 6 (missing parenthesis). In algorithm 1, in the "update posterior" line, I believe this should be $\mathcal{D}_t$. Table 3 is not referred to in section 6.4. ### Significance * I have no concerns regarding significance.

Questions

I have no questions.

Rating

7

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

I do not see any unaddressed limitations.

Reviewer bPdt3/10 · confidence 4/52024-07-12

Summary

The paper proposes a more data-efficient algorithm inspired by Bayesian adaptive experimental design. This algorithm runs maximally informative simulations in a batch-sequential process, estimating posterior distribution parameters and optimal designs by maximizing a variational lower bound of the expected information gain. The algorithm is validated on both synthetic and real-data problems.

Strengths

1. The paper is well-written, offering comprehensive background information and a thorough review of the literature. The method is rigorously compared with other related approaches across multiple metrics on both synthetic and real-data datasets. 2. The paper focuses on a well-motivated and challenging problem.

Weaknesses

The paper can be improved in the following ways: 1. My biggest concern for this paper is its novelty: replacing EIG by a variational lower bound has been explored and well studied by many literature. (e.g. [18], [29]) 2. The method is currently compared only against "Random" and "IMSPE". However, there exist numerous other variants of Bayesian optimal design and frequentist approaches with diverse optimality criteria. It would be valuable to compare the proposed method with these state-of-the-art alternatives. 3. The variational inference is usually applied to sampling from posterior with large-scale datasets or high-dimensional parameter spaces. But the paper only present results on small datasets with low-dimensional parameter spaces, which makes it less convincing.

Questions

1. In Figure 1: (a) It's challenging to ascertain if stability has been achieved, which affects the credibility of the conclusion that BACON achieves rapid convergence in terms of MAP estimates. (b). It's difficult to determine if EIG statistically outperforms others in terms of RMSE. 2. The paper lacks clarity in specifying its contributions and novelty. Could you please elaborate on the main distinction between BACON and other Bayesian optimal design methods that incorporate variational inference (e.g. Variational Bayesian Optimal Experimental Design)? 3. There are numerous instances where variational inference (VI) can fail and produce poor approximations of the target posterior distribution. It would be beneficial to investigate the performance of BACON in scenarios where VI struggles to capture the characteristics of the target distribution, such as when dealing with highly correlated coordinates or multimodal targets.

Rating

3

Confidence

4

Soundness

3

Presentation

3

Contribution

2

Limitations

Variational inference (VI) is typically used to approximate posterior distributions in large-scale datasets and high-dimensional parameter spaces. However, the paper indicates that BACON, which uses VI in Bayesian optimal design, still struggles to scale to large datasets. This raises the question of why VI is used in BACON instead of traditional MCMC sampling methods like HMC. The rationale for choosing VI over these traditional methods remains unclear.

Authorsrebuttal2024-08-14

Thanks for the feedback. We understand the reviewer's concern, as we perhaps have not selected the most appropriate performance metrics to show, and the discussion is lacking a few key insights. We would like to point out that most of the posteriors in the synthetic calibration problem of Sec. 6.2 are multimodal, since the simulators are simply random functions drawn from a GP prior and only 5 "real" data points were provided, alongside 20 initial simulation points. In these multimodal problems, MAP estimates, which show most of the confidence interval overlaps in Fig. 1, should not be considered a primary metric of performance, since the mode of the posterior will often not match the true calibration parameter, given the low amount of data. We, however, have been able to show that our proposed method (BACON) achieves its intended goal, which is to maximise the EIG, as measured by the expected KL divergence from final to initial posterior $\mathbb{D}_{\mathrm{KL}}(p_T||p_0)$ (see Eq. 1 for the equivalence), when compared to the baselines in the paper across all experimental benchmarks. We will clarify these points in the revision and include a few examples of some of the posterior distributions we find in the synthetic and real data problems to better illustrate such challenges.

Reviewer Z11Z7/10 · confidence 3/52024-07-12

Summary

This paper addresses the challenge of calibrating expensive-to-evaluate computer models using Bayesian adaptive experimental design. The novelty of the proposed method (BACON) lies in using the expected information gain (EIG), which is a principled information theoretic criterion for active learning, to perform calibration of models. Another point of departure from the existing literature is the fact that BACON performs active learning in the joint space of design and parameters.

Strengths

This is a technically solid and well-written paper which I really enjoyed reading. The idea of using the EIG criterion over the joint space of $\theta$ and $x$ may sound simple and straightforward once you see it written down that it is almost surprising that no has done this before. Many good papers seem obvious once you read them, and I feel like this paper belongs in that category. Despite being a notation heavy paper, the authors did a good job of making the writing clear and concise.

Weaknesses

I do not have any major concerns. Some questions/comments that might help improve the paper are as follows: * It would have been nice to see an ablation study where the design and parameters are not jointly optimized over, in order to ascertain the benefit of doing so. * The error function $\varepsilon$ is also modellled as a GP. I wonder how does BACON's performance get affected if this assumption does not hold (i.e. when this error model is misspecified). * Perhaps a discussion of the hyperparameters/settings of the proposed method would be nice to have, in terms of how to set them and how sensitive the performance of BACON is to their values. * MMD is not mentioned in the text despite being plotted in Figure 1(d). * Table 3 is not referenced in Section 6.4. * Please include up/down arrows in the tables next to the columns so that it is easy to read them. * Appendix B seems incomplete.

Questions

* Can you say something about the tightness of the bound in Section 5.1? What does it depend on? (I suppose these would have been discussed in ref [13] but it would be nice if mentioned here as well) * Can you explain a bit more why VBMC performs better in the synthetic experiments but fares poorly in the other experiments? Does it have anything to do with the dimensionality of the problem? * Reporting the KL divergence between the prior and the posterior after T iterations tells us how much the posterior has changed, but that does not mean we are converging to the true posterior, right (we may be confidently biased)? Is the KL divergence between $p_T$ and $p^*$ a better indicator for accuracy?

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

The authors have discussed the limitations adequately.

Reviewer bRKR5/10 · confidence 4/52024-07-12

Summary

This paper considers the problem of calibration of computer models as an active learning problem. Given the objective being maximizing the expected information gain (EIG) about calibration parameters, and based on the assumption of linear dependency between simulator outcome and true observation, this work proposes a Gaussian process model that jointly models true observations, calibration parameters and simulator outcomes, and maximization of EIG is performed based on this model. Due to the intractability of the EIG objective, the authors further propose to use variational objective in replace of the original EIG in finding the next design parameters and calibration parameters to sample.

Strengths

- The problem setting is interesting and is indeed important in engineering and physical sciences as computer simulators are often used in those areas. - The approach of using a single GP to model calibration parameters and design parameters jointly is novel.

Weaknesses

- Optimizing calibration parameters and design parameters jointly seems to make the problem harder because dimension of the search space is the sum of both spaces, therefore, the applicability of this method to real-world problems may be limited.

Questions

- The proposed approach certainly makes sense when both spaces of calibration parameters and design parameters are low-dimensional. It may be helpful to include a discussion about choosing appropriate method when the number of calibration parameters is bigger or the number of design parameters is bigger, or both.

Rating

5

Confidence

4

Soundness

2

Presentation

2

Contribution

2

Limitations

Please also include limitation of the dimensionality issue as pointed out above.

Reviewer Z11Z2024-08-09

I thank the authors for their clarifying responses to my questions. The general response outlining the difference between experimental design and Bayesian calibration makes the utility of their approach clearer. I think this paper is a useful contribution when calibrating parameters of a computationally expensive mechanistic model that depends on a design variable (which is slightly different to the kind of models calibrated using simulation-based inference). Hence, I am happy to recommend accept.

Reviewer CjgP2024-08-12

Thanks for the update. I have read the rebuttal and the other reviews and keep my score unchanged.

Reviewer bPdt2024-08-13

Thank you for your response. However, I still find Figure 1 unclear. The confidence intervals of the different methods overlap, suggesting that the proposed method may not be significantly better than the alternatives. As a result, my concerns remain, and I would like to retain my current score.

Reviewer bRKR2024-08-13

Thanks for the response. I have read other comments as well and decide to keep my scoring unchanged.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC