Meta-learning aims at optimizing the hyperparameters of a model class or\ntraining algorithm from the observation of data from a number of related tasks.\nFollowing the setting of Baxter [1], the tasks are assumed to belong to the\nsame task environment, which is defined by a distribution over the space of\ntasks and by per-task data distributions. The statistical properties of the\ntask environment thus dictate the similarity of the tasks. The goal of the\nmeta-learner is to ensure that the hyperparameters obtain a small loss when\napplied for training of a new task sampled from the task environment. The\ndifference between the resulting average loss, known as meta-population loss,\nand the corresponding empirical loss measured on the available data from\nrelated tasks, known as meta-generalization gap, is a measure of the\ngeneralization capability of the meta-learner. In this paper, we present novel\ninformation-theoretic bounds on the average absolute value of the\nmeta-generalization gap. Unlike prior work [2], our bounds explicitly capture\nthe impact of task relatedness, the number of tasks, and the number of data\nsamples per task on the meta-generalization gap. Task similarity is gauged via\nthe Kullback-Leibler (KL) and Jensen-Shannon (JS) divergences. We illustrate\nthe proposed bounds on the example of ridge regression with meta-learned bias.\n
Paper
References (31)
Scroll for more · 19 remaining