In 3D human pose estimation one of the biggest problems is the lack of large,\ndiverse datasets. This is especially true for multi-person 3D pose estimation,\nwhere, to our knowledge, there are only machine generated annotations available\nfor training. To mitigate this issue, we introduce a network that can be\ntrained with additional RGB-D images in a weakly supervised fashion. Due to the\nexistence of cheap sensors, videos with depth maps are widely available, and\nour method can exploit a large, unannotated dataset. Our algorithm is a\nmonocular, multi-person, absolute pose estimator. We evaluate the algorithm on\nseveral benchmarks, showing a consistent improvement in error rates. Also, our\nmodel achieves state-of-the-art results on the MuPoTS-3D dataset by a\nconsiderable margin.\n