Summary
Communicating model uncertainty to human decision makers is critical for trustworthy deployment of assistive AI systems. Prior work, e.g., Vodrahalli et al 2022, have found that people may respond better to miscalibrated AI confidence scores. What underlies this odd behavior? In this submission, the authors explore this phenomenon, specifically developing a new theoretical framework built on structured causal models. In this framework, the authors find that there indeed are some data distributions wherein calibrated AI confidence scores prohibit humans from uncovering optimal policies by which to make their final decision. The authors then propose to resolve this discrepancy through “human-calibration;” multicalibration is demonstrated as one such method. The authors aim to validate their theoretical insights through empirical studies on the same human data from Vodrahalli et al 2022.
Strengths
A substantive assessment of the strengths of the paper, touching on each of the following dimensions: originality, quality, clarity, and significance. We encourage reviewers to be broad in their definitions of originality and significance. For example, originality may arise from a new definition or problem formulation, creative combinations of existing ideas, application to a new domain, or removing limitations from prior results. You can incorporate Markdown and Latex into your review. See /faq.
This paper has many strengths and was an enjoyable read. The paper was very well-written and incredibly well-motivated. As AI systems increasingly move to real-world applications, it is paramount that they are designed with the human user in mind. The theoretical framework from this paper offers a nice contribution to advance our understanding of the nuances of the way that AI output is communicated to humans to best support decision making.
Some further strengths:
- The specific problem of looking at monotonicity vs. discrepancy from monotonicty in confidence values is interesting and nicely expands on prior work.
- The authors make an excellent point on lines 110-114 that human may take diff decisions with same level of confidence; it is good to see that their model incorporates sources of noise ontop of confidence.
- Nice that their method is claimed to be general to other calibration methods.
- The mathematics seems sound and steps naturally follow from previous; though while I’ve attempted to check through it, there is a possibility I have missed something (hence my own confidence score of 3).
Weaknesses
My biggest misgivings about the paper fall into two categories: 1) the discretization of human confidence, and 2) the empirical validation in Section 6. I will discuss each in turn. I am also concerned that the authors do not adequately discuss limitations. But note this further in the Limitations section of the review.
(1) Discretization of human confidence:
- The authors cite two papers from the neuroscience literature suggesting that humans have “discrete confidence”; however, the studies in these papers looks at quite different decision making tasks than are considered here. The tasks are much lower level and rote, e.g., simple perception and motor tasks, at least from my reading. These are quite different from the kinds of rich, high-level cognitive tasks that involve say, coming up with a treatment plan for a patient. As such, I feel the methods here ought to be able to handle finer-grained representations of human uncertainty/confidence rather than too coarse binning.
- I am very open to discussing this — can the authors shed light on how applicable their theoretical results are to different choices of discretization?
- Specifically on Section 6, I have concerns on the choice of 3 bins. 3 bins are incredibly coarse. Why this number? For instance, the “mid” confidence level spans a confidence range of 0.5 (out of the normalized 0-1 scale). I recognize that the binning was done here to ensure approximately equal counts across all bins. However, my sense is that this could yield obfuscate some of the nuance in the results? Have the authors considered a finer-grained decomposition, say to 5 or 7 bins? If nothing else, I think this should be discussed in a limitations sections as I imagine this could be an impactful design choice wrt sensitivity in the empirical results.
(2) Empirical validation
- I am struggling to understand the design of the validation in Section 6. What are the baselines? What are the authors trying to show here to validate their theoretical insights (and which aspects of the theoretical insights?) It is not clear to me what is being shown here. When the authors say “participants benefit from AI advice” in Table 1 [line 333], how is this different from the original paper from which the data is sourced?
- The authors note that the results (which per the above, I’d be keen to hear explicated more – what is the expected behavior of these graphs, relative to what baselines?) are weaker for the Census task. Do the authors have a sense for why this may be?
- The authors grouped the results of cases where participants were told that an AI system had provided the confidence with those told a human had (footnote 9 on pg 8). Is this valid? Could have influenced participants’ uptake of the knowledge? Have the authors explored how results stand up if they decompose / keep those groups separate? Do trends still hold?
Questions
- Many of my questions are listed in the Weaknesses section. In particular, it would be great if the authors could clearly spell out which parts of their experiments (Section 6) validate which parts of their theory. I am open to reconsidering some of my critiques from the Weakness section if this is clarified a bit better.
- Why 8 or 10 bins? [line 298]
- Could you expand please on the confounding factors in the data from other counties? I am not entirely sure why you threw away that data? [lines 287-288] My understanding is this was because those participants were told the AI had a different accuracy? This may be worth exploring (perhaps in future work) – how do your theoretical results hold up across wider swaths of data?)
- How does this relate to RLHF? Please spell out the connection a bit more, or o.w. I actually think it’s unnecessary to mention here. [lines 73-74]
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
The authors do not adequately discuss limitations, nor any potential negative societal impact. I would be keen for the authors to provide their thoughts on these two points during the rebuttal/discussion period.