Summary
This work introduces a novel framework for fODF estimation through equivariant spatio-hemispherical networks that achieve dMRI deconvolution. Experiments on simulated dMRI datasets with known ground truth, as well as on real in vivo dMRI data are conducted, showing promising results while improving over previous methods.
Strengths
1. The paper improves upon previous methods in both processing time, and quantitative results.
2. The evaluation is sound and the experiments nicely show results on synthetic datasets with known ground truth.
Weaknesses
The main weakness of this manuscript is in the way the contributions section is written (at the end of the Introduction section, lines 57--71).
The authors would potentially increase readability of their paper by making this paragraph as clear and as sound as possible.
It would also help readers quickly identify if they wish to continue reading this paper and if it is of interest to their own research.
My suggestions are:
1) Introduce this paragraph by restating what the main aims of this paper are (similar to what you wrote on lines 20--23).
2) Clearly introduce the technical contributions as they are backed by the experiments / results section with a short description of what was achieved.
3) Some of the contributions (specifically, the in vivo qualitative results) are not present in the main manuscript, but are part of the appendix. The exception is Figure 2 which does not have enough description in the main text (lines 254--256), and appears in the middle of a paragraph discussing the synthetic data results. I believe that the experiments section should be clearly reflected in the main aims and contributions of the paper, and the appendix should be used for optional / additional results which do not take away from the main contributions. I understand that there is a limit of 9 content pages to the paper, and I am happy to discuss this further.
Questions
Please find below some questions and general suggestions:
1. Figure 1 introduces the readers to an example of how dMRI data looks, and it is an important prelude towards understanding the problem statement of your paper. For this reason, I suggest the authors include further explanations in this figure, in either visual form or in the captions, such as:
1. How does gradient 25 differ from gradient 288 (maybe try to explain / show that these are different gradient directions and/or strengths instead of the 1/25/288 indices which have no specific meaning in this context)?
2. I think it is also important to show a zoomed-in version of the T1w image, to not confuse the reader that the spatio-spherical signal is also present in the structural data.
2. Please be consistent with referencing your figures in the manuscript: you sometimes write Figure x, and sometimes write Fig. x
3. I suggest you introduce the name of your proposed spatial-hemispherical deconvolution (SHD) framework in the contributions section (lines 57--71) as on the next page Figure 2 shows examples of your model.
4. In Figure 6B, could you discuss whether the low-resolution input to high-resolution output experiments could produce unrealistic reconstructions in the presence of noise / in a real dMRI data setting, as high-angular resolution is needed for higher contrast in the angular domain? I am wondering if for crossing fibers, for example, as shown in Figure 6D, none of the methods can accurately reconstruct the ground truth then maybe we cannot trust these reconstructions for low-angular resolution data?
5. Can you also please label the x-axes in Figures 6A and 6B to make it clear that the values are in degrees?
6. In Figure 6C can you explain the “narrowness” of your result as compared to the ground truth or CSD?
7. Can you provide a short description (in section 4.2.1) of how the peak angular error and false positive rates are calculated? I understand that these are described in A.3, but to improve readability I suggest that they are introduced a bit sooner with the details left for the appendix, or at least to make it more clear that the details are in A.4. Moreover, it would be interesting to understand the slight increase in FPR in both Figures 6A and 6B when comparing SHD(-TV) with RT-ESD.
Limitations
The authors have attempted to address some potential limitations, but would be nice to see a lengthier discussion on how their proposed method would perform on other in vivo clinical datasets, under different noise levels, patient motion, etc.