Strengths
This paper is technically strong, demonstrating a high level of mathematic rigor and diligence. It presents an in-depth description of the proposed estimator. I think the proposed estimator is useful in practice.
Weaknesses
Despite the strong technical details, the paper is poorly written in overall.
__1. Weak motivation__
A current shape of Introduction is weakly motivated. Firstly, the way that the paper motivates the problem is misleading:
> Methods for quantifying the aleatoric uncertainty of the treatment effect have gained surprisingly little attention in the causal machine learning community.
This sentence is incorrect, since there are literatures on quantile regression and semiparametric density estimation, as reviewed in Section 2.
More importantly, the introduction doesn't provide the motivation of the problem against the following question: _why a community need the proposed estimator, given that the quantile estimator can capture the distributional treatment effect_.
__2. Difficult to understand due to insufficient information__
Another issue that the paper has (especially in Introduction) is its lack of back-ground information that readers may need to comprehend. Specifically, in Introduction, the paper doesn’t provide any definition or clue what CDTE is. I understand that the CDTE is defined in the caption of Figure 1 as $P(Y[1] - Y[0] \leq \delta \mid x)$. However, $Y[a]$ is undefined, and this key quantity should be in the main body of the text. Also, even if Figure 1 aims to provide a whole summary of the paper, authors at Introduction have insufficient knowledge to comprehend it. In other words, Figure 1 is too detailed to be presented in Introduction section. Since Figure 1 can only be understood by those who entirely digested the paper from the beginning to the end, Introduction section is not the right position where the Figure 1 is located. I understand the goal of Figure 1, but it doesn't achieve the goal because of insufficient background information. The same issue happens to Figure 2 and Table 1. For example, in Table 1, technical terms like AIPTW, hold-out residual, optimization assumptions are undefined, so it's hard to appreciate the contribution of this paper.
__3. Fuzzy description on contribution__
First, the terms like "Aleatoric uncertainty" and "distributional treatment effect" are used for denoting the same target quantity. Given that the "distributional treatment effect" is clearly describing the problem, I don't see why the authors want to use "Aleatoric uncertainty" as a title and employ these two words to denote the same estimand.
Second, the contribution is wrongly described. Consider this sentence:
> AU-learner solves all of the above-mentioned challenges 1 – 3 .
The AU-learner doesn't address Challenge 1 and 2. Challenge 1 means that the distributional treatment effect is not identifiable. The Makarov's bound, _NOT_ AU-learner, is employed to address Challenge 1. Challenge 2 is actually the same as Challenge 1, since it means that there are no known nuisances-based functional for the distributional treatment effect. Again, the Makarov's bound is used to address Challenge 2. AU-learners are representing the approximated quantity of the target estimand in terms of nuisance functionals.
Finally, third contribution "flexible deep learning instantiation of our AU-learner" is scarcely described only in Section 5. The description needs to be much improved.
__4. Little focus on the real contribution__
The real contribution of this paper, compared to the existing works in Table 1, is to provide the doubly robust conditional distributional treatment effect for the bounds of the conditional CDF of treatment effects. However, little focus and efforts have been made for this contribution. For example, if developing a doubly robust estimator is a contribution, then corresponding results such as detailed error analysis, a closed form of estimators, a detailed recipe of the proposed estimator for the specified working model, how to minimize the losses in Equations (8,9), a simple example, and assumptions should be described.
Questions
1. Are rate-doubly-robustness and Neyman orthogonality violated when the scaling hyperpameters are not $1$?
2. Is the Makarov bound sharp?
3. What are the practical examples of the distributional treatment effect?
4. In line 178, is this phenomenon officially termed the selection bias?