Despite the recent advances in video classification, progress in\nspatio-temporal action recognition has lagged behind. A major contributing\nfactor has been the prohibitive cost of annotating videos frame-by-frame. In\nthis paper, we present a spatio-temporal action recognition model that is\ntrained with only video-level labels, which are significantly easier to\nannotate. Our method leverages per-frame person detectors which have been\ntrained on large image datasets within a Multiple Instance Learning framework.\nWe show how we can apply our method in cases where the standard Multiple\nInstance Learning assumption, that each bag contains at least one instance with\nthe specified label, is invalid using a novel probabilistic variant of MIL\nwhere we estimate the uncertainty of each prediction. Furthermore, we report\nthe first weakly-supervised results on the AVA dataset and state-of-the-art\nresults among weakly-supervised methods on UCF101-24.\n