Summary
The paper considers the causal effect estimation setting where a confounder is latent but with unstructured text data that could serve as proxies. Specifically, the paper proposes to incorporate zero-shot classifiers (to operate on text-based proxies), together with a falsification heuristic, into the proximal causal inference framework. The goal of the text-based proxy design is to craft W and Z that satisfy the identification condition in proximal causal inference framework. Empirical evaluations are presented in synthetic and semi-synthetic experiments.
---
**Post rebuttal**
I have increased my score, under the assumption that the revised manuscript could sufficiently address the original concerns on material organizations.
Strengths
The strength of the paper comes from the attempt to conduct proximal causal inference on unstructured text data. The paper carefully considers identification conditions specified in the proximal causal inference framework, and proposes a design procedure together with a falsification heuristic to find two text-based proxies (so that the proximal inference can be conducted).
Weaknesses
The weakness of the paper comes from the organization of the material (especially the constraints/conditions involved), and relatively simple settings considered in (fully-/semi-) synthetic experiments. In particular, further clarifications and/or discussions on following points would be very helpful (detailed in "Questions" section):
(1) regarding a series of different constraints, conditions, gotcha's, assumptions
(2) the fully and semi- synthetic experiments consider structured data, it is not exactly clear how these text proxies look like and if they correspond to all or part of the aforementioned constraints/conditions. Examples of text proxies would be much more intuitive than claiming "satisfy by design" (line 136)
Questions
(1) regarding a series of different constraints, conditions, gotcha's, assumptions
In Section 2, there are P1 -- P4, a set of conditions should be satisfied in order for the proximal g-formula to work. In order to respond to the criticism of potential unavailability of W and Z in structured data, further assumptions S1 -- S2 are proposed. Then in Section 3 and Section 4, a set of gotcha's, two additional pre-conditions (lines 216 -- 218), and another set of conditions related to odds-ratio based heuristics (lines 245 -- 247) are presented. How do these conditions fit together to respond to the criticism of proximal causal inference on structured data? Are they all identification conditions for the text-based proxy proximal causal inference (other than the fact that they are conditions specified by previous works for different components put together)?
(2) the fully and semi- synthetic experiments consider structured data, it is not exactly clear how these text proxies look like and if they correspond to all or part of the aforementioned constraints/conditions
For fully and semi- synthetic experiments, the setting is largely structured data instead of text-based proxies. Examples of text proxies that the proposed design procedure yields would be much more intuitive than just claiming "satisfy by design" (line 136).
Limitations
The organization of material could be improved to better present how the proposed approach responds to criticism of proximal causal inference (with structured data). Examples beyond fully and semi- synthetic experiments will make it clearer w.r.t. how the designed text-based proxies help to address the aforementioned criticism on previous work.