We study the problem of learning to estimate the 3D object pose from a few\nlabelled examples and a collection of unlabelled data. Our main contribution is\na learning framework, neural view synthesis and matching, that can transfer the\n3D pose annotation from the labelled to unlabelled images reliably, despite\nunseen 3D views and nuisance variations such as the object shape, texture,\nillumination or scene context. In our approach, objects are represented as 3D\ncuboid meshes composed of feature vectors at each mesh vertex. The model is\ninitialized from a few labelled images and is subsequently used to synthesize\nfeature representations of unseen 3D views. The synthesized views are matched\nwith the feature representations of unlabelled images to generate pseudo-labels\nof the 3D pose. The pseudo-labelled data is, in turn, used to train the feature\nextractor such that the features at each mesh vertex are more invariant across\nvarying 3D views of the object. Our model is trained in an EM-type manner\nalternating between increasing the 3D pose invariance of the feature extractor\nand annotating unlabelled data through neural view synthesis and matching. We\ndemonstrate the effectiveness of the proposed semi-supervised learning\nframework for 3D pose estimation on the PASCAL3D+ and KITTI datasets. We find\nthat our approach outperforms all baselines by a wide margin, particularly in\nan extreme few-shot setting where only 7 annotated images are given.\nRemarkably, we observe that our model also achieves an exceptional robustness\nin out-of-distribution scenarios that involve partial occlusion.\n