Response on the remaining concerns
We sincerely appreciate your detailed feedback and the time you've taken to review our manuscript.
> I was not aware of the resemblance to [16] pointed out by reviewer xhDC, which changes my assessment on novelty and therefore less prone to raise the score further.
We'd like to address the perceived resemblance between our work and [16] and clarify our contributions, which we partly discussed in the 'Related work' section.
While we acknowledge [16] as an inspiring and original contribution to the field, we believe that our work also contributes distinctly and significantly.
Similarities:
* The use of the neighbor-joining (NJ) algorithm that maps continuous coordinates to a phylogenetic tree
* Both works focus on Bayesian phylogenetic inference
Differences:
* We developed a variational inference (VI) algorithm, while [16] developed an MCMC-based algorithm.
* We addressed the issue of parameterizing $Q(\tau)$ in the VI algorithm to cover a vast number of tree topologies, while [16] highlighted the fidelity of hyperbolic spaces to embed the distribution over trees with distances $P(\tau, B_\tau | Y)$.
* For the use of the NJ, we explicitly defined a distribution over topologies $Q(\tau) = E_{Q(z)}[I[\tau(z) = \tau]]$ instead of mapping coordinates $z$ to $(\tau, B_\tau)$ as seen in [16]. This distinction is crucial as our approach avoids the issue of the Jacobian determinant seen in [16].
* Unlike [16], our link function $\tau(z)$ does not necessarily rely on NJ, as we don't directly link $z$ to $B_\tau$. We showcase the results using UPGMA in Fig. R3 Fourth.
Distinct contributions not comparable to [16]:
* We introduced a tractable lower bound $\mathcal{L}$, then explored designs of variational distributions and control variates to complete a novel VI algorithm (GeoPhy).
* We benchmarked the model evidence estimations (MLLs) across approaches and exhibited significant improvement over other methods that considered whole topologies.
We hope these clarifications address your concerns.
> the paper would benefit from mentioning the faster learning runtime of VCSMC and Vaiphy ...
Thank you for your suggestion. Accordingly, we will include a discussion on the trade-off of the performance and runtimes between these methods in their standard use in our revised manuscript.
> R2 Fourth: I think that the DS1 dataset is not preferable for this analysis, as the support of the MrBayes posterior includes few topologies.
> Datasets DS4 and higher would've been a better selection
Thank you for the valuable suggestions. In response, we've investigated DS4 and DS7 alongside DS1 to discuss the limitations of $Q(\tau)$, especially in cases with more diffused tree samples.
Also, we present the difference of tree topology distributions more clearly by showing (b) the frequency of the most frequent topology and (c) the number of topologies up to their cumulative frequency matches to 95 percentile, in addition to (a) the diversity index.
| DS1 | MrBayes | VBPI-GNN | GeoPhy |
|--|--:|--:|--:|
| (a) Simpson's diversity index | 0.87 | 0.86 | 0.36 |
| (b) Top freq. topology | 0.27 | 0.26 | 0.79 |
| (c) #topology up to 95% freq. | 42 | 44 | 11 |
| DS4 | MrBayes | GeoPhy |
|--|--:|--:|
| (a) | 0.90 | 0.68 |
| (b) | 0.28 | 0.55 |
| (c) | 208 | 58 |
| DS7 | MrBayes | GeoPhy |
|--|--:|--:|
| (a) | 0.99 | 0.99 |
| (b) | 0.02 | 0.02 |
| (c) | 753 | 553 |
For DS4, the overall tendencies in GeoPhy: lower (a), higher (b), and lower (c), are observed as DS1. Interestingly, for DS7, we observed that GeoPhy also represents more diverse tree samples than DS1 and DS4. However, the number of unique topologies up to 95% freq. is still lower than MrBayes, which implies the requirement of more expressiveness on $Q(\tau)$ to represent fine topology weights.
While the results for VBPI-GNN are not readily available for DS4 and DS7 due to time constraints, we would like to include the corresponding results in our revised manuscript.
> Furthermore, what in R2 First-Fourth informs the diversity of the trees sampled by $Q(\tau)$?
Given that most tree topologies $\tau$ had very low frequencies, we used the bipartition frequencies of species defined for each tree topology edge as a more concise statistic for the distribution of $Q(\tau)$.
In Fig. R2, we present bipartitions ordered by descending frequency as seen in MrBayes. The intermediate values between 0 and 1 in this frequency plot reveal topology diversities, where VBPI-GNN aligns more closely with MrBayes for the DS1 dataset compared to GeoPhy. As GeoPhy tends to take values near zero and one in the slope region, it indicates a need for increased expressiveness of $Q(\tau)$ to better represent intermediate frequencies. This trend was also noticeable for DS4. For DS7, while GeoPhy traced the curve more accurately, fluctuations around this curve highlight potential room for improvement. We intend to incorporate these figures in our updated manuscript.