We present an approach for unsupervised learning of speech representation\ndisentangling contents and styles. Our model consists of: (1) a local encoder\nthat captures per-frame information; (2) a global encoder that captures\nper-utterance information; and (3) a conditional decoder that reconstructs\nspeech given local and global latent variables. Our experiments show that (1)\nthe local latent variables encode speech contents, as reconstructed speech can\nbe recognized by ASR with low word error rates (WER), even with a different\nglobal encoding; (2) the global latent variables encode speaker style, as\nreconstructed speech shares speaker identity with the source utterance of the\nglobal encoding. Additionally, we demonstrate an useful application from our\npre-trained model, where we can train a speaker recognition model from the\nglobal latent variables and achieve high accuracy by fine-tuning with as few\ndata as one label per speaker.\n