This paper focuses on weakly-supervised action alignment, where only the\nordered sequence of video-level actions is available for training. We propose a\nnovel Duration Network, which captures a short temporal window of the video and\nlearns to predict the remaining duration of a given action at any point in time\nwith a level of granularity based on the type of that action. Further, we\nintroduce a Segment-Level Beam Search to obtain the best alignment, that\nmaximizes our posterior probability. Segment-Level Beam Search efficiently\naligns actions by considering only a selected set of frames that have more\nconfident predictions. The experimental results show that our alignments for\nlong videos are more robust than existing models. Moreover, the proposed method\nachieves state of the art results in certain cases on the popular Breakfast and\nHollywood Extended datasets.\n