Human pose estimation from single images is a challenging problem in computer\nvision that requires large amounts of labeled training data to be solved\naccurately. Unfortunately, for many human activities (\\eg outdoor sports) such\ntraining data does not exist and is hard or even impossible to acquire with\ntraditional motion capture systems. We propose a self-supervised approach that\nlearns a single image 3D pose estimator from unlabeled multi-view data. To this\nend, we exploit multi-view consistency constraints to disentangle the observed\n2D pose into the underlying 3D pose and camera rotation. In contrast to most\nexisting methods, we do not require calibrated cameras and can therefore learn\nfrom moving cameras. Nevertheless, in the case of a static camera setup, we\npresent an optional extension to include constant relative camera rotations\nover multiple views into our framework. Key to the success are new, unbiased\nreconstruction objectives that mix information across views and training\nsamples. The proposed approach is evaluated on two benchmark datasets\n(Human3.6M and MPII-INF-3DHP) and on the in-the-wild SkiPose dataset.\n