Semi-supervised learning objectives as log-likelihoods in a generative model of data curation
We currently do not have an understanding of semi-supervised learning (SSL)\nobjectives such as pseudo-labelling and entropy minimization as\nlog-likelihoods, which precludes the development of e.g. Bayesian SSL. Here, we\nnote that benchmark image datasets such as CIFAR-10 are carefully curated, and\nwe formulate SSL objectives as a log-likelihood in a generative model of data\ncuration that was initially developed to explain the cold-posterior effect\n(Aitchison 2020). SSL methods, from entropy minimization and pseudo-labelling,\nto state-of-the-art techniques similar to FixMatch can be understood as\nlower-bounds on our principled log-likelihood. We are thus able to give a\nproof-of-principle for Bayesian SSL on toy data. Finally, our theory suggests\nthat SSL is effective in part due to the statistical patterns induced by data\ncuration. This provides an explanation of past results which show SSL performs\nbetter on clean datasets without any "out of distribution" examples. Confirming\nthese results we find that SSL gave much larger performance improvements on\ncurated than on uncurated data, using matched curated and uncurated datasets\nbased on Galaxy Zoo 2.\n
Paper
References (38)
Scroll for more · 26 remaining