Monocular 3D object detection has recently shown promising results, however\nthere remain challenging problems. One of those is the lack of invariance to\ndifferent camera intrinsic parameters, which can be observed across different\n3D object datasets. Little effort has been made to exploit the combination of\nheterogeneous 3D object datasets. In contrast to general intuition, we show\nthat more data does not automatically guarantee a better performance, but\nrather, methods need to have a degree of 'camera independence' in order to\nbenefit from large and heterogeneous training data. In this paper we propose a\ncategory-level pose estimation method based on instance segmentation, using\ncamera independent geometric reasoning to cope with the varying camera\nviewpoints and intrinsics of different datasets. Every pixel of an instance\npredicts the object dimensions, the 3D object reference points projected in 2D\nimage space and, optionally, the local viewing angle. Camera intrinsics are\nonly used outside of the learned network to lift the predicted 2D reference\npoints to 3D. We surpass camera independent methods on the challenging KITTI3D\nbenchmark and show the key benefits compared to camera dependent methods.\n