Weaknesses
First of all, I think that the presentation style of the theoretical part of this paper leaves much to be desired: (i) very often notation is used before it is introduced; moreover, in many cases this notation is used without being introduced/defined anywhere in the main part of the paper; (ii) sometimes the notation is inconsistent; (iii) lot's of typos, and I do not mean just textual ones, there are typos in the formulae in the main part of the paper. These issues make it really hard to understand and evaluate theoretical contribution of the paper. Below I list some examples of such issues occurring only in Section 2 and subsection 2.1 (it's just two and a half pages):
1. In lines 068-069 $Y$ stands for classification tasks, in lines 070-071 $Y$ is referred to as labels corresponding to classification tasks.
2. Content random variables $Z$ come from distribution with density $q_{Z}$ in many places in the paper (e.g. line 068), but in line 182 (Assumption 2, item 5) the content density is denoted by $p_{Z}$.
3. Speech features $X\sim q_{\gamma}$ in line 064 and speaker identity $G\sim\gamma$ in line 069. Neither $\gamma$ nor $q_{\gamma}$ are defined. If $\gamma$ is distribution of the random variable $G$ corresponding to speaker identity, then $q_{\gamma}$ should be its density. Why then do speech features $X$ come from the distribution with the density $q_{\gamma}$?
4. Line 077 (Assumption 1, item 3) $d$ is not defined. Functions $I$ and $h$ are not defined either. I understand that they denote mutual information and entropy, but it's better to state that explicitly, especially if you also use the same notation $h$ in line 267 for a multi-class classifier.
5. It's better to be precise and call $B_{t}$ a Wiener process rather than just some Markov process (line 086-087), because later you explicitly use that diffusion process (2) is the one relying on Gaussian noise when you write down Formula (5).
6. In Formula (2) there should be just $X_{t}$ instead of $X_{t}^{\leftarrow}$.
7. Lines 089-090: word "wehe".
8. Formula (3): there should have been $f(X_{t}^{\leftarrow}, A, T-t)$ instead of $f(X_{t}^{\leftarrow}, A, t)$ in the drift coeficient of the reverse diffusion similarly to diffusion coefficient $\nu(T-t)$. The same mistake in the drift coefficient can be found in Formula (7) and the one in lines 147-150, while Formula (10) has the right drift term $f(X_{t}^{2\leftarrow 1}, A, T-t)$.
9. In the same Formula (3) there must be $dB_{t}^{\leftarrow}$ instead of $dB_{t}$ since we switch to the reverse-time dynamics.
10. In lines 097-098 $\chi$ is not defined.
11. In lines 121 we start considering speech features $X^{1}\sim q_{\beta,0}$ while before we had speech features with densities $q_{\gamma}$ and $q_{\alpha}$. None of $\alpha$, $\beta$ or $\gamma$ are defined in the paper up to this moment. Am I correct that $\gamma$ are all possible speech features, and $\alpha$ and $\beta$ correspond to speech features for two particular speakers? Moreover, $\gamma$ is used later in the definition of the implicit bottleneck variable (Formula (11)), and it certainly has another meaning than when it is used in the expressions like $q_{\gamma}$.
12. In Formula (9) the drift term should be $f(X_{t}^{1},A,t)dt$ rather than just $-f(X_{t}^{1},A,t)$ without time differential.
13. In Formula (10) $\hat{\theta}$ is not defined.
14. In Formula (10) you use $B_{t}^{\leftarrow}$ which must be Wiener process in reverse-time, but it is never mentioned in the paper.
15. The acronym AEVC in lines 127-128 is not defined. Does it mean "Autoencoder Voice Conversion"?
16. In Formula (11) in the definition of the implicit bottleneck variable (which is key to the whole theoretical framework you've designed) you use undefined notation $G_{<t}$. This indeed makes it very difficult to understand what is going on after in this section.
17. In the formula in lines 147-150 in the third term you should have $d\tau$ instead of $dt$.
18. In Formula (13) an undefined function $\hat{g}$ is used.
19. In Definition 3 you use $d_{TV}$ distance without mentioning what it is.
20. In Assumption 2, item 1 you define $t^*$, but do not use it anywhere in this Assumption.
There is also a mistake in Formula (5): it suggests that the score function is $-(X_{t}-X_{0})/\sigma^{2}(t)$ for all $t$, but this is not true. E.g. in case of Variance Preserving diffusions [1] the score function is $-(X_{t}-a(t)X_{0})/\sigma^{2}(t)$ where $X_{t}|X_{0}\sim a(t)X_{0}+\sigma(t)\mathcal{N}(0,1)$ and $a(t)$ varies in $t$.
The next weakness regards the notion of implicit bottleneck variable. First of all, I feel that the definition itself is ambiguous in the sense that the functions $\eta$, $\zeta$ and $\gamma$, if they exist, are not unique, and they are not even unique up to some linear transform that could preserve "information" contained in, for example, $\zeta(X_{t}^{\leftarrow})$. A simple example is as follows: suppose $X_{t}^{\leftarrow}=|X_{0}^{\leftarrow}|$, then we can have either $\zeta(X_{t}^{\leftarrow})=X_{0}^{\leftarrow}$ and $\eta(z,g)=|z|$, or $\zeta(X_{t}^{\leftarrow})=|X_{0}^{\leftarrow}|$ and $\eta(z,g)=z$. And further assumptions (Assumption 2) rely heavily on properties of $\zeta^{\star}=\zeta(\hat{X}_{<t^{*}}^{\leftarrow})$, e.g. on values like $I(\zeta^{\star};X)$ that are different in the two cases.
The former issue leads to the next concern: we really do not know much about this implicit bottleneck variable $\zeta^{\star}$. What are the sufficient conditions for it to exist? Are items 1 and 4 in the Assumption 2 feasible? This issue is very important, because you build your theory in Section 2 based on the random variable $\zeta^{\star}$ which is not even guaranteed to exist (it corresponds to the case $t^{*}=0$ if I'm not mistaken). Apart from this, I also have some less important considerations which I'll list below as questions.
As for experimental part, I also feel it has serious drawbacks. As far as I understand, the whole theoretical framework derived in Section 2 is tested only on synthetic datasets while the authors claim that they describe the mechanism behind disentanglement in general diffusion models. Moreover, after this claim they start describing their setup in "intuitive" voice conversion terms which may be a bit confusing. If they do so, it would be good to demonstrate some of their findings on at least one example not related to voice conversion.
In the experiments on synthetic datasets, the authors consider linear subspace diffusion models which decompose speech features following Equation (18). The papers you refer to before this equation to justify this decomposition deal with discriminative tasks, but you consider voice conversion, a generative task. I don't know papers that successfully use this type of decomposition in generative tasks, and it seems very unlikely that such a simplified point of view on speech features decomposition can lead to a good quality of generated speech. So, I do not think that the results of these experiments on synthetic datasets can be extrapolated to real voice conversion.
The experiments on realistic datasets correspond to theoretical findings in Section 3 which are actually more related to ensembling theory than to disentanglement in diffusion models. Altough I do not have much experience in ensembling methods, I think that the first conclusion the authors make ("Adding target speakers reduces speaker distortion") is not surprising and can be explained not only by Theorem 4, but by the reduced variance as well: when the number of target speakers for voice conversion increases, the variance introduced by different performance of classifiers on different (imperfectly generated) voices gets lower resulting in improved performance in classification task when some kind of aggregation mechanism is used. Two other conclusions ("Different VC excels at different tasks" and "Tradeoff between classifier accuracy and diversity") are quite interesting, but they can not be explained by the theoretical framework the authors have developed.
Overall, I think that the theory trying to explain disentanglement in diffusion models presented in this paper lacks convincing experimental material, has some issues regarding feasibility of the assumptions crucial for this theory, and the text itself has to be made more readable by getting rid of numerous typos/mistakes/inconsistencies in notation.
[1] Score-Based Generative Modeling through Stochastic Differential Equations, Song et al.
Questions
1. What is the conceptual difference between $A$ and $Z_{t}$ in Formula (4)? They both seem to correspond only to content information. When you define $A$, you refer to [2], but in that paper the score function depended only on $X_{t}$, $A$, $t$ and speaker information, there was not additional bottleneck $Z_{t}=z_{\phi, t}(X_{t})$ instead of $X_{t}$. Also, does this formula mean that $z_{\phi}$ is a separate neural network? Is it supposed to be trained just jointly with the score-matching network $s_{\theta}$?
2. In lines 128 - 130, why do you call $X_{T}$ a time-dependent variable in contrast with time-independent $A$? $X_{T}$ is also independent of diffusion time $t$ and depends only on a fixed time horizon $T$.
3. I do not understand Formula (13). If $a=1$ and $b=2$, then $\hat{X}^{a\to b}$ should be the result of forward and reverse diffusions defined in Equations (9-10), but I cannot see how the Definition 1 of the implicit bottleneck and (9-10) imply Formula (13). What also bothers me is that you define the implicit bottleneck variable by considering the critical timestep $t^{*}$ for the reverse process $X_{t}^{\leftarrow}$, and it is the diffusion time for reverse-time processes. But in Formula (13) you use expressions like $\zeta(X_{<t^{\star}}^{a})$, but here time $t$ bounded by the critical $t^{\star}$ already corresponds to forward-time process. Is it correct? To increase presentation clarity, I would recommend to use different notation for forward and reverse timesteps (e.g. $t$ and $\tau$).
4. Regarding Formula (13), I also don't understand what is $\Theta^{a\to b}$. You only say that it is "random noise introduced by the noising process". Also, what happens with this formula if $t^{*}=0$? Is in this case $\hat{X}^{a\to b}$ a function only of this "random noise"?
5. In Definition 3, do you really mean $max$ rather than $sup$? What happens if the domain of $Z$ is unbounded?
6. Item 2 in Assumption 2 suggests that $\hat{G}$ and $G$ have the same dimensionality. But $G$ is some abstract "speaker identity random variable" while $\hat{G}$ is a speaker embedding from some speaker verification system whose dimensionality can be different. What happens if dimensionalities $\hat{G}$ and $G$ do not match?
[2] Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme, Popov et al.