We present an approach for identifying the most walkable direction for\nnavigation using a hand-held camera. Our approach extracts semantically rich\ncontextual information from the scene using a custom encoder-decoder\narchitecture for semantic segmentation and models the spatial and temporal\nbehavior of objects in the scene using a spatio-temporal graph. The system\nlearns to minimize a cost function over the spatial and temporal object\nattributes to identify the most walkable direction. We construct a new\nannotated navigation dataset collected using a hand-held mobile camera in an\nunconstrained outdoor environment, which includes challenging settings such as\nhighly dynamic scenes, occlusion between objects, and distortions. Our system\nachieves an accuracy of 84% on predicting a safe direction. We also show that\nour custom segmentation network is both fast and accurate, achieving mIOU (mean\nintersection over union) scores of 81 and 44.7 on the PASCAL VOC and the PASCAL\nContext datasets, respectively, while running at about 21 frames per second.\n