Deep reinforcement learning (RL) uses model-free techniques to optimize\ntask-specific control policies. Despite having emerged as a promising approach\nfor complex problems, RL is still hard to use reliably for real-world\napplications. Apart from challenges such as precise reward function tuning,\ninaccurate sensing and actuation, and non-deterministic response, existing RL\nmethods do not guarantee behavior within required safety constraints that are\ncrucial for real robot scenarios. In this regard, we introduce guided\nconstrained policy optimization (GCPO), an RL framework based upon our\nimplementation of constrained proximal policy optimization (CPPO) for tracking\nbase velocity commands while following the defined constraints. We also\nintroduce schemes which encourage state recovery into constrained regions in\ncase of constraint violations. We present experimental results of our training\nmethod and test it on the real ANYmal quadruped robot. We compare our approach\nagainst the unconstrained RL method and show that guided constrained RL offers\nfaster convergence close to the desired optimum resulting in an optimal, yet\nphysically feasible, robotic control behavior without the need for precise\nreward function tuning.\n