Thank You for Your Constructive Feedback and Suggestions
Dear Reviewers,
We sincerely appreciate your time and effort in reviewing our submission. Your insights are invaluable, and we are grateful for the constructive feedback provided throughout this process. Your comments have significantly helped us refine and clarify our contributions.
> ## 1. The Paper’s Main Idea
This paper proposes reorienting Inverse Reinforcement Learning (IRL)-based Imitation Learning (IL) from data alignment to task alignment. We have worked to clearly articulate the core idea of our work, as outlined in our response:
* The **goal** of our method is to learn a policy that fulfills tasks as formulated in Definition 1, which is general and encompasses many real-world examples, as demonstrated in [1] and as proved in our individual response.
* We introduced the concept of task-aligned reward functions and explained our **insight** that achieving high utility under these task-aligned reward functions is essential for task fulfillment.
* Our **approach** treats expert demonstrations as weak supervision to learn a set of reward functions that capture task-aligned rewards and then trains a policy to achieve high utility across this set.
> ## 2. Complexity and Relevance of Benchmarks
We have emphasized that the MiniGrid environment is a well-established RL benchmark with its own publication at NeurIPS 2023 [2]. Despite its seemingly simple appearance, MiniGrid presents significant challenges for current IRL/IL algorithms due to its **massive state space** (partially observable, random layout in every episode, etc.), and other complexities. We will further clarify this in our revised manuscript to address any concerns regarding the complexity of the environments we used.
> ## 3. Expanded Experimental Evaluation
In response to requests for more diverse and challenging experiments, we have included new experimental results, such as **continuous control tasks** and additional baselines like **f-IRL** and **RECOIL** [3]. These results demonstrate the broader applicability of our method in both online and **offline IL** and the effectiveness of our method across different types of environments, further reinforcing the robustness of our approach.
> ## 4. Moving Forward
We have carefully considered your feedback and will take several steps to further improve our submission. We plan to enhance the clarity of our contributions and streamline the theoretical sections to ensure that the main ideas are communicated more effectively. Additionally, we will expand our experimental results on other complex benchmarks, such as Locomujoco.
Once again, we sincerely thank you for your thoughtful reviews and contributions to improving our submission. We are confident that the revisions made in response to your feedback will result in a stronger and more impactful paper.
Best regards,
The Authors
[1] Abel, David, et al. On the expressivity of Markov reward, NeurIPS 2021
[2] Chevalier-Boisvert, Maxime, et al. “Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks.” NeurIPS 2023
[3] Sikchi et al.; Dual RL: Unification and new methods for reinforcement and imitation learning, ICLR 2024