Human Expertise in Algorithmic Prediction

We introduce a novel framework for incorporating human expertise into algorithmic predictions. Our approach leverages human judgment to distinguish inputs which are algorithmically indistinguishable, or "look the same" to predictive algorithms. We argue that this framing clarifies the problem of human-AI collaboration in prediction tasks, as experts often form judgments by drawing on information which is not encoded in an algorithm's training data. Algorithmic indistinguishability yields a natural test for assessing whether experts incorporate this kind of "side information", and further provides a simple but principled method for selectively incorporating human feedback into algorithmic predictions. We show that this method provably improves the performance of any feasible algorithmic predictor and precisely quantify this improvement. We find empirically that although algorithms often outperform their human counterparts on average, human judgment can improve algorithmic predictions on specific instances (which can be identified ex-ante). In an X-ray classification task, we find that this subset constitutes nearly $30\%$ of the patient population. Our approach provides a natural way of uncovering this heterogeneity and thus enabling effective human-AI collaboration.

Paper

Similar papers

Peer review

Reviewer XFzN8/10 · confidence 4/52024-06-25

Summary

This paper introduces a new framework into algorithmic predictions. The paper asks and answers the question "how can we incorporate human input into the prediction algorithm, which may not even be captured in the training data"? The authors develop a method that first runs the predictor, and then runs a second predictor using the human input. The authors show that even a simple instantiation of their method can outperform existing predictors. They use the X-ray classification task as experimental datasets.

Strengths

The paper is written very clearly, and offers a novel method to incorporate human input into algorithmic prediction. Both theoretical derivations and experiment results are sound. The contributions of this paper is significant, and I believe this paper deserves to be accepted in its current form.

Weaknesses

The paper would be even more satisfying if the method is presented as a framework rather than a specific instantiation. In addition, it would be great if the authors can discuss potential ways to improve on the method they propose, and what these methods mean in the broader context of incorporating human feedback into algorithmic predictions. Nevertheless, these small weaknesses does not diminish the significance and novelty of this paper.

Questions

My main comment is that the authors should comment more about the future work and implications of this method. Furthermore, I would be interested to hear what the authors think about a related paper [1], and how these papers might be related. [1] DEFINING EXPERTISE: APPLICATIONS TO TREATMENT EFFECT ESTIMATION (https://arxiv.org/pdf/2403.00694)

Rating

8

Confidence

4

Soundness

4

Presentation

4

Contribution

4

Limitations

The authors have addressed the limitations in the conclusion section

Reviewer WEMo7/10 · confidence 4/52024-07-07

Summary

The paper proposes a framework to incorporate human expert knowledge in algorithmic predictions. Under this framework, the authors introduce a meta-algorithm that uses a training dataset including human expert predictions together with a multi calibrated partition of the data; a partition of the dataset into bins, where each bin contains data that are indistinguishable to the predictive model. Using the data of each bin the meta-algorithm trains a regression algorithm to predict the true label from the human expert prediction. In this way, the authors aim to leverage the human expertise, that may be more accurate than the predictive algorithm on specific instances, to achieve complimentary—to achieve higher predictive accuracy through human AI collaboration than the performance of a human expert or AI in isolation.

Strengths

The paper suggests an elegant method to improve algorithmic predictions in light of human expertise, that could have significant applications such as the medical domain, where the additional information of human experts may lead them to more accurate predictions on certain instances compared ot predictive models. The paper is very well and clearly written, nicely motivated and follows a clear structure. There is a thorough and comprehensive discussion on related work as well as a comprehensive and clearly presented experimental evaluation.

Weaknesses

Since the theoretical results of section 6 complement the ones of section 4, it would be perhaps more natural to follow them, rather than placing them after the experimental evaluation, which appears a bit odd.

Questions

N/A

Rating

7

Confidence

4

Soundness

3

Presentation

4

Contribution

3

Limitations

The authors adequately discuss the limitations of their work.

Reviewer jnGw7/10 · confidence 3/52024-07-13

Summary

The paper first presents some theory for the modelling of how to identify when human judgements may offer a better diagnosis - through access to additional information - than machine predictions, despite the latter typically being more accurate. This is followed by exploring how to integrate the human input with the algorithmic (model) input. Subsequently, the authors present some focussed experimental results using chest x-ray interpretation that support their proposition.

Strengths

Originality: carefully drawn comparison with the literature, situates and differentiates the contribution. Quality + Clarity (addressed together): Clear abstract and intro with well-defined contributions. Content offers a reasonable balance between technical and intuitive. Recognition of the value of the human contribution and seeking to integrate it in decision making. The later mathematical results (section 4) have effective accompanying interpretations (see complementary point in weaknesses). Effective, selective presentation of results: choosing one and going into detail, while two other cases in the appendices support the same observation, rather than trying to squeeze them all into the paper body. Same applies to results in section 5.2. Significance: provides a sound framework for a particular, amenable class of collaboration problems that allows for the proper incorporation of human prediction where machine prediction could fall short.

Weaknesses

Clarity: Indistinguishability and multicalibration are critical elements to the contribution; it would be helpful if the interpretation of their definitions (3.1, 3.2) went into a bit more detail for accessibility. This reader is not succeeding in following the argument about robustness (section 6).

Questions

Q1. The case studies are retrospective so both machine and human outcomes are available to use in the analysis. How would the approach work in a live situation?

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Section 7 provides some properly reflective critique on scope and applicability.

Reviewer fseb8/10 · confidence 3/52024-07-18

Summary

This paper introduces a framework for joint human-AI prediction, where human experts can augment AI predictions in particular ex ante identifiable subsets.

Strengths

This paper makes a lot of interesting contributions. First, its scope is broad and important: it tackles the question of how and whether human judgment can improve the predictions of any learning algorithm. That is and will remain to be a very important question in our time. It contributes a very interesting framework, rooted in algorithmic indistinguishability and multicalibration, to find subsets in which no algorithm in a user-specified class has predictive power (because they are algorithmically indistinguishable) but human experts do (because they might have more access to the instances, such as doctors examining patients). It demonstrates that using this framework, we can find subsets of instances where human experts can outperform algorithms, and thus the combination of the two can outperform either alone. It applies this to an important medical problem and in another domain of making predictions from photos of people. It even extends the framework to apply to a setting with noncompliance. The community stands to learn a lot from this paper.

Weaknesses

As the authors mention, the framework is dependent on minimizing mean squared error only.

Questions

How might you model deicision makers with richer preferences than mean squared error?

Rating

8

Confidence

3

Soundness

4

Presentation

4

Contribution

4

Limitations

Yes.

Reviewer WEMo2024-08-09

I would like to thank the authors for their reply. The suggested changes by the authors address my point and should be done to improve the flow.

Reviewer XFzN2024-08-11

I recommend acceptance

Many thanks for the authors' detailed response. I am happy to see that all reviewers unanimously recommended acceptance. Therefore, I am happy to accept the paper, and nominate the paper for awards if the AC agrees.

Reviewer jnGw2024-08-13

Thanks for the detailed and helpful response both to me and the other reviewers.

Feiyu Zhu22024-12-19

Hello, thank you for your interesting paper, but as for your example of Doctor A and Doctor B, I think it is the reverse causality. In your example, Doctor A, except for patients with hypertension, usually follows the advice of the algorithm. This situation is often because the algorithm has more errors in the judgment of patients with hypertension, which is not as accurate as the doctor's own judgment. At other times, the judgment is accurate, and the doctor does not need to invest too much effort and judgment. You deduce from this that the optimal algorithm is minimizes error on patients who do not have high blood pressure. I think this judgment is irresponsible. It is difficult for me to understand how the example you give contributes to the explanation of Section VI

Rohan Alur12024-12-19

Thank you for your comment. The example is correct as written — the physician only uses the algorithm’s recommendations on a subset of the distribution, and our goal is to minimize error on this subset. Our results are agnostic as to why this is the case; in particular, we do not interrogate whether the physician is “correct” in their judgment. Furthermore, these kinds of restrictions can arise for other reasons — for example, it may be that a certain physician (or hospital) is not equipped to see patients with a certain condition, in which case a similar restriction arises. We take this conditioning event as given, and seek to minimize error over the corresponding conditional distribution.

Program Chairsdecision2024-09-25

Decision

Accept (oral)

© 2026 NYSGPT2525 LLC