One of the major challenges for autonomous vehicles in urban environments is\nto understand and predict other road users' actions, in particular, pedestrians\nat the point of crossing. The common approach to solving this problem is to use\nthe motion history of the agents to predict their future trajectories. However,\npedestrians exhibit highly variable actions most of which cannot be understood\nwithout visual observation of the pedestrians themselves and their\nsurroundings. To this end, we propose a solution for the problem of pedestrian\naction anticipation at the point of crossing. Our approach uses a novel stacked\nRNN architecture in which information collected from various sources, both\nscene dynamics and visual features, is gradually fused into the network at\ndifferent levels of processing. We show, via extensive empirical evaluations,\nthat the proposed algorithm achieves a higher prediction accuracy compared to\nalternative recurrent network architectures. We conduct experiments to\ninvestigate the impact of the length of observation, time to event and types of\nfeatures on the performance of the proposed method. Finally, we demonstrate how\ndifferent data fusion strategies impact prediction accuracy.\n