In this paper we consider a problem known as meta-learning, consisting of building policies that achieve good generalization performance and adapt quickly to different tasks. In our novel formulation, which we denote by cross-learning, we introduce meta-learning through the coupled optimization of a set of rewards that are defined for different tasks. This coupling is effected by a projection step that brings the task-specific policies close to a central one which combines the information collected across tasks. Since such a projection is computationally expensive, we derive a relaxed version that can be obtained in closed-form through a geometric rule. While our initial cross-learning formulation is widely general, and connects with state-of-the art strategies, we specialize it for the case of reinforcement learning. In particular, we search for continuous policies in reproducing kernel Hilbert spaces, which are considered in order to avoid discretization and bypass parametric models. Preliminary numerical experiments performed on the classical cartpole system corroborate that the cross-learned control policy performs well in different scenarios.
Paper
Full text
Meta-Learning through Coupled Optimization in Reproducing Kernel Hilbert Spaces
Semantic Scholar · Computer Science · 2019
Abstract
In this paper we consider a problem known as meta-learning, consisting of building policies that achieve good generalization performance and adapt quickly to different tasks. In our novel formulation, which we denote by cross-learning, we introduce meta-learning through the coupled optimization of a set of rewards that are defined for different tasks. This coupling is effected by a projection step that brings the task-specific policies close to a central one which combines the information collected across tasks. Since such a projection is computationally expensive, we derive a relaxed version that can be obtained in closed-form through a geometric rule. While our initial cross-learning formulation is widely general, and connects with state-of-the art strategies, we specialize it for the case of reinforcement learning. In particular, we search for continuous policies in reproducing kernel Hilbert spaces, which are considered in order to avoid discretization and bypass parametric models. Preliminary numerical experiments performed on the classical cartpole system corroborate that the cross-learned control policy performs well in different scenarios.