Coping with the variability in humans reward during simulated human-robot interactions through the coordination of multiple learning strategies
An important current challenge in Human-Robot Interaction (HRI) is to enable\nrobots to learn on-the-fly from human feedback. However, humans show a great\nvariability in the way they reward robots. We propose to address this issue by\nenabling the robot to combine different learning strategies, namely model-based\n(MB) and model-free (MF) reinforcement learning. We simulate two HRI scenarios:\na simple task where the human congratulates the robot for putting the right\ncubes in the right boxes, and a more complicated version of this task where\ncubes have to be placed in a specific order. We show that our existing MB-MF\ncoordination algorithm previously tested in robot navigation works well here\nwithout retuning parameters. It leads to the maximal performance while\nproducing the same minimal computational cost as MF alone. Moreover, the\nalgorithm gives a robust performance no matter the variability of the simulated\nhuman feedback, while each strategy alone is impacted by this variability.\nOverall, the results suggest a promising way to promote robot learning\nflexibility when facing variable human feedback.\n