Recent works have shown that generative data augmentation, where synthetic\nsamples generated from deep generative models complement the training dataset,\nbenefit NLP tasks. In this work, we extend this approach to the task of dialog\nstate tracking for goal-oriented dialogs. Due to the inherent hierarchical\nstructure of goal-oriented dialogs over utterances and related annotations, the\ndeep generative model must be capable of capturing the coherence among\ndifferent hierarchies and types of dialog features. We propose the Variational\nHierarchical Dialog Autoencoder (VHDA) for modeling the complete aspects of\ngoal-oriented dialogs, including linguistic features and underlying structured\nannotations, namely speaker information, dialog acts, and goals. The proposed\narchitecture is designed to model each aspect of goal-oriented dialogs using\ninter-connected latent variables and learns to generate coherent goal-oriented\ndialogs from the latent spaces. To overcome training issues that arise from\ntraining complex variational models, we propose appropriate training\nstrategies. Experiments on various dialog datasets show that our model improves\nthe downstream dialog trackers' robustness via generative data augmentation. We\nalso discover additional benefits of our unified approach to modeling\ngoal-oriented dialogs: dialog response generation and user simulation, where\nour model outperforms previous strong baselines.\n