We present a framework to translate between 2D image views and 3D object\nshapes. Recent progress in deep learning enabled us to learn structure-aware\nrepresentations from a scene. However, the existing literature assumes that\npairs of images and 3D shapes are available for training in full supervision.\nIn this paper, we propose SIST, a Self-supervised Image to Shape Translation\nframework that fulfills three tasks: (i) reconstructing the 3D shape from a\nsingle image; (ii) learning disentangled representations for shape, appearance\nand viewpoint; and (iii) generating a realistic RGB image from these\nindependent factors. In contrast to the existing approaches, our method does\nnot require image-shape pairs for training. Instead, it uses unpaired image and\nshape datasets from the same object class and jointly trains image generator\nand shape reconstruction networks. Our translation method achieves promising\nresults, comparable in quantitative and qualitative terms to the\nstate-of-the-art achieved by fully-supervised methods.\n