In this paper, we describe an approach to achieve dynamic legged locomotion\non physical robots which combines existing methods for control with\nreinforcement learning. Specifically, our goal is a control hierarchy in which\nhighest-level behaviors are planned through reduced-order models, which\ndescribe the fundamental physics of legged locomotion, and lower level\ncontrollers utilize a learned policy that can bridge the gap between the\nidealized, simple model and the complex, full order robot. The high-level\nplanner can use a model of the environment and be task specific, while the\nlow-level learned controller can execute a wide range of motions so that it\napplies to many different tasks. In this letter we describe this learned\ndynamic walking controller and show that a range of walking motions from\nreduced-order models can be used as the command and primary training signal for\nlearned policies. The resulting policies do not attempt to naively track the\nmotion (as a traditional trajectory tracking controller would) but instead\nbalance immediate motion tracking with long term stability. The resulting\ncontroller is demonstrated on a human scale, unconstrained, untethered bipedal\nrobot at speeds up to 1.2 m/s. This letter builds the foundation of a generic,\ndynamic learned walking controller that can be applied to many different tasks.\n