We consider the task of generating diverse and novel videos from a single\nvideo sample. Recently, new hierarchical patch-GAN based approaches were\nproposed for generating diverse images, given only a single sample at training\ntime. Moving to videos, these approaches fail to generate diverse samples, and\noften collapse into generating samples similar to the training video. We\nintroduce a novel patch-based variational autoencoder (VAE) which allows for a\nmuch greater diversity in generation. Using this tool, a new hierarchical video\ngeneration scheme is constructed: at coarse scales, our patch-VAE is employed,\nensuring samples are of high diversity. Subsequently, at finer scales, a\npatch-GAN renders the fine details, resulting in high quality videos. Our\nexperiments show that the proposed method produces diverse samples in both the\nimage domain, and the more challenging video domain.\n