Summary
The paper investigated the persistent homology with multiple filtration parameters. The authors proposed a framework, T-CDR, which generalized the previous studies in multi-parameter persistent homology. They further presented stability and convergence guarantees on S-CDR, which is a special case of T-CDR introduced to ensure robustness. Third, their theoretical claim and contributions were supported by the empirical convergence studies as well as classification tasks on several immunohistochemistry datasets.
Strengths
1. [Originality] The proposed framework, T-CDR in Definition 1 not only generalized previous work in candidate decomposition (e.g., MPI) as well as approaches in rank-invariant (MPL, MPK). I believe that it opens up new research on ensuring different guarantees by varying parameters/operators in (1).
1. [Significance] The stability and convergence guarantee (Theorem 1 and (8)) of S-CDR, which is a special case of T-CDR, provides a more stable and useful multi-parameter persistent homology framework.
1. [Quality] The authors supplement the convergence rate claim with informative empirical convergence studies, showing a clear trend of error rate reduces with $~n^{-1/2}$ as $n$ grows matching eq. (8).
Weaknesses
1. In the empirical convergence rate experiment, it was not clear to me what is the “ground truth” representation you are comparing against. For synthetic data in Figure 3, it makes sense that you can get access to the density $f$ (therefore $\mathcal F_{C, f}$ and $\mathbb M$) given that you generated the data from some probability distribution. How do you get the $\mathbb M$ for the immunohistochemistry data (as in Figure 4)? Provide more clarifications on this will be beneficial.
1. It seems like MPL runtime can be improved 25x-50x with the sparse implementation in Algorithm 4 (see, e.g., Row #2 of Table 2 vs. Rows #4 and #6 of Table 2); this suggests that the runtime win might be due implementation rather than faster algorithmic time complexity. The claim will be more convincing if the authors can provide more insights, justifications, or analyses of the time complexity as to why the proposed algorithm is more efficient than prior work.
1. What is the bifiltration parameters used for each experiments in Sections 4? Are they radius and co-density for Figure 3, and CD8 and CD68 for the immunohistochemistry data? Adding more explanation there will increase the clarity more.
1. In Sections 1-3, the function $f$ is used to define the multi-parameter filtration function; however, in Section 4, the notation is defined as the density (e.g., in L254). I would suggest to choose another notation to avoid confusion.
Questions
1. It looks to me most of Section 4.1 is a continuation of Theorem 1, is there any specific reason to put this in this section rather than in Section 3 (and make it a Proposition/Corollary)?
1. Empirically, how sensitive is the algorithm for the larger intrinsic dimension $d$? If I understand the experiments correctly, they all seem to have $d=2.$ I am curious about how the intrinsic dimension $d$ will impact convergence and/or performance.
1. I am also curious about how the proposed framework can be applied in higher-order homology descriptors (empirically). This is also related to Question #2 above.
1. TDA has been applied in numerous different domains such as galaxy, proteins, single-cell sequencing, 3D CAD point clouds, medical imaging [A-C] etc. I am curious if the proposed multi-parameters filtraion work can be expanded in fields outside immunohistochemistry.
1. How optimal/tight is the bound in (8)? Can we get a better convergence result if we choose op, $\omega$, and/or $\phi$ differently?
1. [Minor] Given that T-CDR is a generalization of both the rank-invraint and candidate decomposition method, should the name template “candidate decomposition” representation be modified to better reflect what it can be capable of?
1. [Minor/Typo?] Should the citation in L477 be [CB20] instead of the PersLay paper?
---
[A] Wasserman, Larry. “Topological Data Analysis.” Annual Review of Statistics and Its Application 5 (2018): 501–32.
[B] Chen, Yu-Chia, and Marina Meila. “The Decomposition of the Higher-Order Homology Embedding Constructed from the k-Laplacian.” Advances in Neural Information Processing Systems 34 (2021).
[C] Wu, Pengxiang, Chao Chen, Yusu Wang, Shaoting Zhang, Changhe Yuan, Zhen Qian, Dimitris Metaxas, and Leon Axel. “Optimal Topological Cycles and Their Application in Cardiac Trabeculae Restoration.” In International Conference on Information Processing in Medical Imaging, 80–92. Springer, 2017.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
Authors have addressed the limitations of their work. Negative social impact is not applicable because this work is a theoretical contribution.