We thank the reviewer for the useful and constructive comments. We answer the additional questions as follows.
> Supplement to Q1: As I mentioned earlier, the example of Car-Circle seems somewhat simplistic. If time permits, could you please add the generation results of OASIS to Figure 8c?
- Figure 8c shows the CVAE reconstruction results for the Car-Circle task, supplementing Figure 7c. During the rebuttal phase, we added the OASIS generation results to Figure 8c as requested by the suggestions:
> "_Q1: Compared with CVAE, OASIS shows that the conditional diffusion model has a better ability to generate according to the condition information, as shown in Figure 7c, but this example is a bit simple. Can the author show more comparisons of the two similar to Figure 7c? For example, add the OASIS generation results to Figure 8c?_"
- The Figure R-1 we provided in the rebuttal phase is the similar generation results of OASIS to those of CVAE shown in Figure 8c. To make a clear visualization, we (1) reduce the number of trajectories; (2) add different conditions for generation; and (3) visualize the generation results from OASIS and CVAE in separate figures. The random seeds to sample trajectories are different, so the targets for reconstruction are slightly different in Figure R-1 and Figure 8c. However, we keep the same set of trajectories for reconstruction for OASIS and CVAE to make a clear and fair comparison in Figure R-1.
- We select the Car-Circle task for this visualization experiment because (1) it is widely used in offline safe RL benchmark [1] and related offline safe RL works [2, 3, 4]; (2) its state space contains the position information of ego agent, which is easy to visualize in 2D space.
- For the request to visualize other tasks, we will update the visualization results for other robots (i.e., Drone) that have high-dimensional observation and action space and complicated dynamics models. Since we can not update the PDF file at this stage, we will include these in the appendix of the revised manuscript.
> Supplement to Q3: Yes, I agree that the conclusion you provided is more accurate and reasonable.
We thank the reviewer for the agreement and acknowledgment. We have added related discussions in our revised manuscript.
---
[1] Zuxin Liu, et al. "Datasets and benchmarks for offline safe reinforcement learning." arXiv preprint arXiv:2306.09303 (2023).
[2] Yinan Zheng, et al. "Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model." ICLR 2024
[3] Zijian Guo, et al. "Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning." ICML 2024.
[4] Kihyuk Hong, et al. "A primal-dual-critic algorithm for offline constrained reinforcement learning." AISTATS 2024.