Segmenting objects in videos is a fundamental computer vision task. The\ncurrent deep learning based paradigm offers a powerful, but data-hungry\nsolution. However, current datasets are limited by the cost and human effort of\nannotating object masks in videos. This effectively limits the performance and\ngeneralization capabilities of existing video segmentation methods. To address\nthis issue, we explore weaker form of bounding box annotations.\n We introduce a method for generating segmentation masks from per-frame\nbounding box annotations in videos. To this end, we propose a spatio-temporal\naggregation module that effectively mines consistencies in the object and\nbackground appearance across multiple frames. We use our resulting accurate\nmasks for weakly supervised training of video object segmentation (VOS)\nnetworks. We generate segmentation masks for large scale tracking datasets,\nusing only their bounding box annotations. The additional data provides\nsubstantially better generalization performance leading to state-of-the-art\nresults in both the VOS and more challenging tracking domain.\n