Strengths
The main strength of the paper is a novel method for designing randomized experiments under network interference. The novelty of the method comes from developing an objective function whose decision variables are continuous, (essentially) representing the covariance matrix of assignments. This stands in contrast to previous methods which focus on selecting clusters for independent cluster randomization designs. An additional strength of the paper is that the objective considers both bias and variance, which is uncommon: most approaches attempt to "fix" one of these two things, but typically not both.
One of the most exciting technical connections is the use of Grothendieck's identity, which (to the best of my knowledge) has not been used in causal inference. This offers a new tool in the design of experiments, which I suspect will be useful beyond the specific design used in this paper.
Moreover, the simulations are very well formulated and executed.
In particular, the authors investigate the effectiveness of their design under various forms of model misspecification, which speaks to the robustness of the design.
This is important to investigate via simulations, as formal theory seems difficult given the black-box nature of the experimental design.
Weaknesses
There are several weaknesses of the current method.
1. **Assumed Knowledge**: The paper makes some strong assumptions as to what the experimenter knows. For example, authors assume that the coefficients $\alpha_i$ are known by the experimenter. I think most experimenters will find that knowledge of each individual $\alpha_i$ is too strong of an assumption to be practical -- in the SUTVA setting, this implies that all individual treatment effects can be estimated perfectly without any randomization. In order to be transparent, authors should state this assumption and estimator earlier in the paper, perhaps replacing the HT estimator in (eq 5).
2. **Pre-specified Clusters**: The paper assumes that clusters are pre-specified and little advice is given to practicioners on how to select the clusters in order to minimize MSE.
3. **Understanding of the Variance**: In order to run power calculations, experimenters should have at least a rough understanding (e.g. asymptotic rates) of how the variance of the experimental design depends on sample size $n$ and network parameters. Because this method is based on a black-box optimization procedure, it seems hard to analyze the variance (i.e. optimal value).
4. **Confidence Intervals**: In practice, experimenters value interval estimators (i.e. confidence intervals) more than point estimators, as they provide a method for uncertainty quantification. The necessary tools for uncertainty quantification (e.g. Central Limit Theorem, variance estimation) are not presented in this paper.
While authors should make certain minor changes to address the above, I think that these weaknesses actually constitute further research directions on this exciting method.
Overall, it is my opinion that the strengths of the paper outweigh the weaknesses.
In the two sections below, I discuss minor weakenesses in the technical discussions and literature review which should be addressed by authors before publication.
## Technical Remarks
Below are some minor remarks on technical aspects of the paper.
I think there are a few technical issues in the discussions, but I believe these can be easily fixed and will strengthen the technical contribution of the paper.
1. (Line 137) authors write "without loss of generaltiy, we consider the balanced cluster-level randomization scheme satisfying...$E[z_i] = 1/2$." I think that this restriction is perfectly fine, but I would say it is not technically correct to describe it as "without loss of generality". Consider the usual SUTVA setting: if the outcomes under treatment have more variation than outcomes under control, then the Horvitz--Thompson estimator can be made to have smaller variance by setting $p$ so that treatment is assigned more frequently. I think the phrase "without loss of generality" is not warranted and a simple fix is to just remove it.
2. (Line 152) The word "overparametrized" is not quite standard in this literature so I'd recommend either informally defining it, or just saying that "there are more unknown potential outcomes than observations". This is true even under SUTVA.
3. (Line 208): Authors introduce what they call "Assumption 2". I would refer to this as a "condition" rather than an assumption. The reason is that Assumption 2 only plays a role in choosing the design -- using standard techniques (CLT + variance estimator), Assumption 2 would not be necessary for, say, the validity of confidence intervals. Some readers might misinterpret the use of the word "Assumption" to think "if this condition does not hold, then the estimates are no longer statistically valid in some sense".
4. (Line 254) Authors write "This lemma enables us to sample from bivariates Bernoulli distribution with mean (1/2, 1/2) and any valid covariance". While this is true for $n=2$ variables, I do not believe it to be true for general $n$ variables. It is sort of implied in the Section that this "Grotendeick mapping" can recover any covariance matrix of $\pm 1$ variables. If authors have a proof of this, they should provide it; otherwise, they should clearly state that this "Groethendeick mapping" cannot generate all $\pm 1$ covariance matrices. I think that clarifying this will increase understanding and appreciation of the method.
5. (Line 294): The Monte Carlo simulation is only performed 200 times. In my experience, this is quite low. Can you increase to 10,000 Monte Carlo runs before camera ready submission? This will increase the reader's confidence in you results.
## Literature Review + References
The authors have missed a few important references in the causal inference literature.
I believe these should be easily fixable and would increase the relevancy of the paper by better situating it within the causal inference literature.
1. (Line 36-37) In reference to cluster designs, authors write: "This technique is originally developed in [31] and becomes a prevalent paradigm for network experiment design". I completely agree that [31] was an influential paper for introducing cluster designs to the computer science community. [35] works the exposure mapping framework for interference [1] which allows for this arbitrary network interfernece. However, the so-called "partial interference" assumption has been used since at least Hudgens & Halloran (2008) and cluster designs were advocated for here. So, I'd at least reference some of this early work on clustering in the context of partial interference. I think this will also tie your contributions back to the vien of causal inference in the statistics literature in a stronger way.
2. (Line 38) Authors write that "sharing same treatment within cluster is usually necessary for characterizing the GATE". I would remove or substantially weaken this statement. The proliferation of cluster designs is *not* because they are necessary; but rather, a conceptually simple yet effective type of experimental design.
3. (Line 64) Authors write "In this paper, we propose to treat the covariance matrix of treatment vector as a decision variable in optimization". The following paper seems especially relevant and authors should draw a comparison: Harshaw et al (2019) "Balancing Covariates in Randomized Experiments using the Gram--Schmdit Walk Design". This paper studies experimental designs that directly control the covariance matrix Cov(z) (via discrepancy theory) to bound the variance of Horvitz--Thompson estimator by an implicit ridge regression of outcomes on covariates. A key idea in that paper is to use the operator norm as a measure of worst-case variance, which seems like an alternative to your Assumption 2. Given the similarity in the spirit of the two papers, this paper would benefit from a brief comparison discussion.
4. (Line 94) Authors cite several papers on bipartite experiments. It seems that [15] and [16] are duplicates. To the best of my knowledge, the paper of Zigler & Papadogeorgou (2021) was the first paper to propose bipartite experiments (it has been a working paper since 2019), so a citation is warranted in that discussion.
5. (Line 104): Authors write "Along the same direction, [32] tries to provide..." I recommend revising this language. The word "tries" gives the indication that "[32] tries and fails".
6. (Line 108): In the discussion of partial interference, early work like Hudgens and Halloran (2008) is missing.
7. (Line 116): A citation of several papers with more general forms of interference is listed. The recent paper Harshaw, Sävje, Wang (2022) "A design-based riesz representation framework for randomized experiments" is worth citing, as it proposes a deisgn-based framework which captures and extends previous types of interference.
8. (Line 251): I haven't read [19], but it was my understanding that Grothendieck's identity is typically a different type of statement, where there are two fixed vectors $x$ and $y$ and the random variables are $\textrm{sign}(\langle z , x \rangle)$ and $\textrm{sign}(\langle z , y \rangle)$, where $z$ is uniform from the $\ell_2$ ball. In fact, I have only seen Lemma 1 in certain course notes on Sums-of-Squares (though I'm sure it's appeared in other places). If Lemma 1 does not directly appear in [19] then authors should cite a relevant paper that derives it. In fact, this would probably be helpful to tie your work back to theoretical computer science's use of the technique.
## References
- Harshaw, C., Sävje, F., Spielman, D., & Zhang, P. (2019). "Balancing covariates in randomized experiments with the Gram-Schmidt Walk design". (arXiv:1911.03071)
- Harshaw, C., Sävje, F., & Wang, Y. (2022). "A design-based riesz representation framework for randomized experiments". (arXiv:2210.08698)
- Hudgens, M. G., & Halloran, M. E. (2008). "Toward causal inference with interference". Journal of the American Statistical Association, 103(482), 832–842.
- Zigler, C. M. and Papadogeorgou, G. (2021). "Bipartite causal inference with interference". Statist. Sci., 36(1):109–123.