In this work, we present a learning-based pipeline to realise local\nnavigation with a quadrupedal robot in cluttered environments with static and\ndynamic obstacles. Given high-level navigation commands, the robot is able to\nsafely locomote to a target location based on frames from a depth camera\nwithout any explicit mapping of the environment. First, the sequence of images\nand the current trajectory of the camera are fused to form a model of the world\nusing state representation learning. The output of this lightweight module is\nthen directly fed into a target-reaching and obstacle-avoiding policy trained\nwith reinforcement learning. We show that decoupling the pipeline into these\ncomponents results in a sample efficient policy learning stage that can be\nfully trained in simulation in just a dozen minutes. The key part is the state\nrepresentation, which is trained to not only estimate the hidden state of the\nworld in an unsupervised fashion, but also helps bridging the reality gap,\nenabling successful sim-to-real transfer. In our experiments with the\nquadrupedal robot ANYmal in simulation and in reality, we show that our system\ncan handle noisy depth images, avoid dynamic obstacles unseen during training,\nand is endowed with local spatial awareness.\n