Summary
The paper considers a (finite-horizon) dynamic persuasion problem
between a sender and a receiver, both of whom are long-lived. At each
time $t$, there is a publicly-observable state $s_t$ and a
payoff-relevant quantity $\theta_t$, which is observed only by the
sender, and whose distribution depends on the current state (and
is independent of other quantities). Based on the observation of
$\theta_t$ the sender recommends an action to the receiver (which may
or may not be followed). The publicly-observable state then updates to
$s_{t+1}$ according to a transition kernel that depends on the current
state $s_t$, the quantity $\theta_t$ and the action chosen by the
receiver. Both the sender and the receiver seek to maximize their
total expected payoffs.
Previous work on this topic, with few exceptions, has focused on
myopic receivers, motivated by settings in which the receiver is
short-lived. The difference here then is the focus on long-lived,
far-sighted receiver. In this setting, the paper first shows (via an
example) that the class of Markovian signaling schemes (whose
recommendations only depend on the current state $s_t$) is
insufficient for optimal persuasion, and the sender can do better
using a signaling scheme that takes into account the history of the
process. Due to the computational difficulties in working with general
history-dependent schemes, the paper then considers promise-form
signaling schemes, which make recommendations based not only on the
current state, but also on a (history-dependent) "promise", which is a
guarantee on the receiver's continuation payoffs. Essentially, the
promise succinctly summarizes the history, thereby reducing the
computational complexity to be polynomial in the size of set of
promises. The authors show that, upon imposing a honesty condition on
the promises across time, the class of promise-form signaling schemes
suffice for optimal persuasion. The authors also propose an
approximation scheme for computing approximately-persuasive
promise-form signaling schemes with good payoff guarantees, that is
polynomial in the approximation factor.
Strengths
+ The paper considers an interesting variation of the sequential
persuasion problem, allowing for far-sighted receivers. This makes
the problem substantially more complex. Nevertheless, the paper
identifies a class of relatively simple and approximately persuasive
signaling schemes that nevertheless achieve optimal payoffs for the
sender, and are furthermore computationally tractable.
+ The paper illustrates well the insufficiency of the class of
Markovian signaling schemes, and furthermore (adapting existing
results) shows that finding a constant-factor approximation within
the class of Markovian signaling schemes is NP-hard.
+ The class of promise-form signaling schemes is fairly simple and
easy to implement; furthermore, it seems approximately-optimal such
schemes can be computed via solving an LP (repeatedly).
Weaknesses
+ While the class of promise-form signaling schemes is interesting,
there is a significant line of work in economics that studies the
use of promises in repeated games with incomplete information. The
paper does not cite those papers, nor does it place its
contributions within that context. A particularly relevant paper in
this line is Abreu, Pierce and Stacchetti (Econometrica, 1990),
whose results imply the sufficiency of the class of "promise-form"
strategies for discounted repeated games with imperfect monitoring.
+ Similarly, the paper would benefit from connecting with general
literature on repeated games (with or without incomplete
information). For instance, the insufficiency of Markov signaling
schemes is very much in the same vein as the inefficiency of Markov
perfect equilibria in, say, repeated prisoner's dilemma to sustain
cooperation. With far-sighted receivers, it is not surprising that
Markov signaling schemes are not optimal for the sender.
+ $\epsilon$-persuasiveness: In the analysis of history-dependent (or
promise-form) signaling schemes, the authors relax the
persuasiveness requirement to $\epsilon$-persuasiveness. There is a
subtle issue in interpreting this relaxation. To explain, a natural
relaxation would be that the receiver's expected continuation payoff
from following the recommendation is at most $\epsilon$ worse than
choosing any other action, *after* receiving the recommendation.
Specifically, the expectation taken here would be with respect to
the posterior belief after receiving the recommendation. However,
the condition in Definition 1 requires something different; it
states that the receiver's expected payoff from following a
recommendation, *multiplied* by the probability of receiving that
recommendation, should be at most $\epsilon$ worse. In particular,
there is an extra factor equaling the probability of recommending a
particular action.
While this may seem like a minor technical issue, this has
substantial implication on the assumption that the receiver would
adopt such a recommendation. For instance, this suggests that as
long as the probability of recommending an action is small, the
sender can recommend an action that can yield substantially lower
continuation payoff for the receiver, and still expect the receiver
to accept the recommendation. This seems to be a very strong
assumption on the receiver's behavior, that does not align with the
assumption that the receivers are (approximately) Bayesian.
Moreover, with such a strong assumption, it is no longer clear if
$OPT$ is the right benchmark for comparison.
A potential fix to this issue would be to impose the relaxation on the
conditional expectation, i.e., to replace the $\epsilon$ term in the
definition with $\epsilon \sum_{\theta}
\mu_h(\theta|s_h)\phi_\tau(a|\theta)$. However, it is not clear if
the later approximation results continue to apply with this change.
+ Finally, while the paper makes sound and rigorous technical
contribution, there is not enough discussion motivating the specific
model being studied. For instance, there is no discussion of the
motivation behind far-sightedness assumption; the myopic behavior of
the receivers in previous work is frequently motivated by assuming a
series of short-lived receivers. In particular, are there any
specific applications where a single sender and a single receiver
interact in the manner studied? (I think this is especially useful
given the somewhat complicated form of the signaling scheme
proposed.) Some discussion here would benefit the paper by grounding
the theoretical results.
Questions
+ Do the approximation results continue to hold if the relaxation of
the persuasiveness constraint is imposed on the conditional
expectation?
+ With the current definition of $\epsilon$-persuasiveness, it may be
possible to design mechanisms that achieve payoffs substantially
better than $OPT$. Are there any guarantees on how small (or large)
this difference can be?
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Limitations
The assumptions are stated clearly. Some discussion of the limitations
induced by relaxing the persuasiveness/honesty requirements would be
helpful.