Summary
The paper studies the problem of estimating potential outcomes in the presence of a combinatorial number of intervention choices. Under some assumptions, they propose a two-phased algorithm "Synthetic Combinations": first exploit structure across combinations of interventions (via "horizontal regression") and then exploit structure across units (via "vertical regression"). Experiments are given in the appendix.
Strengths
The proposed algorithm is clean and intuitive. It also seems to scale nicely with the number of intervention combinations in experiments.
While the theoretical guarantees rely heavily on a bunch of assumptions, Section 7 proposes an experimental design framework which ensures that an important set of assumptions (existence of donor units) will be met with high probability. In fact, I strongly propose that the authors rephrase their paper to highlight this; otherwise it is hard to believe that their work will useable as it is highly unlikely that all the required assumptions are met in practice without having control in assigning interventions to the units.
Weaknesses
I did not check all the proofs in detail, but I do not see any glaring weaknesses.
There are a lot of assumptions and it is highly unlikely that all the required assumptions are met in practice (Section 7 helps to mitigate some of these concerns).
I am skeptical about the low-rank assumption on the matrix of Fourier coefficients $A$. While it is true low-rank assumptions are common in prior matrix completion settings, they usually directly consider the matrix at hand and not transform it into the Fourier space first. For example, Lines 655-657 in the appendix writes "This missingness pattern where outcomes with larger absolute values are observed is common in applications such as recommendation engines, where we are only likely to observe ratings for combinations that users either strongly like or dislike". The corresponding missingness pattern in the problem studied here is on the $N$-by-$2^p$ matrix. It is unclear to me why it should be believable that the transformed space is low-rank. The authors ought to justify this, ideally with practical examples/settings, or risk diminishing the impact of their contributions.
Questions
Line 83:
By "equivalent", do you mean that they proved equivalence between the two problems via reductions, or do you mean "equivalent" in a colloquial sense of the word?
Assumption 3.1:
As discussed in the weaknesses, it is unclear to me why this model is interesting or justified. Of course, this work can be appreciated under the restriction of this assumption, but it will greatly weaken the contributions. I am more than happy to increase my "contribution" score if the authors provide sufficient justification for the low-rank assumption.
Type on Line 183:
double "exists"
Motivating example on Line 190:
I don't understand why this motivates the existence of donor units when the paper has thus far repeatedly claim to allow unobserved confounding. If we allow interventions to be arbitrarily assigned to units, it is unclear why we should believe that donor units exist. The "correct" way to justify should be to say that there is an experimental design that ensures the existence of donor units, and then refer to Section 7.
Determining donor set on Line 246:
This feels very ad-hoc. As it is unlikely that donor units will exist if we allow arbitrary experiments, I feel that this paragraph could be removed once the authors reorder their paper to place more emphasis on the experimental design proposed in Section 7.
Subsection on Additional Assumptions:
I feel that "so-and-so also has such an assumption" is not sufficient discussion of assumptions. Firstly, "so-and-so" may have the assumptions under different contexts (e.g. see my complaint about low-rank assumption in the Weaknesses section) so it is unclear why such assumption is justified in the setting studied in this paper. Secondly, the discussion should explain "what goes wrong" if one particular assumption is violated, or why we should expect any particular assumption to hold in practice. As mentioned several times by now, one "partial fix" is to emphasize that experimental design of Section 7 guarantees some assumptions with high probability. That is, "Synthetic Control" should be used in conjunction with the experimental design proposed in Section 7.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.