In this paper, we propose an end-to-end trainable regression approach for\nhuman pose estimation from still images. We use the proposed Soft-argmax\nfunction to convert feature maps directly to joint coordinates, resulting in a\nfully differentiable framework. Our method is able to learn heat maps\nrepresentations indirectly, without additional steps of artificial ground truth\ngeneration. Consequently, contextual information can be included to the pose\npredictions in a seamless way. We evaluated our method on two very challenging\ndatasets, the Leeds Sports Poses (LSP) and the MPII Human Pose datasets,\nreaching the best performance among all the existing regression methods and\ncomparable results to the state-of-the-art detection based approaches.\n