As data volumes continue to grow, the labelling process increasingly becomes\na bottleneck, creating demand for methods that leverage information from\nunlabelled data. Impressive results have been achieved in semi-supervised\nlearning (SSL) for image classification, nearing fully supervised performance,\nwith only a fraction of the data labelled. In this work, we propose a\nprobabilistically principled general approach to SSL that considers the\ndistribution over label predictions, for labels of different complexity, from\n"one-hot" vectors to binary vectors and images. Our method regularises an\nunderlying supervised model, using a normalising flow that learns the posterior\ndistribution over predictions for labelled data, to serve as a prior over the\npredictions on unlabelled data. We demonstrate the general applicability of\nthis approach on a range of computer vision tasks with varying output\ncomplexity: classification, attribute prediction and image-to-image\ntranslation.\n