Emotional voice conversion (EVC) aims to convert the emotion of speech from\none state to another while preserving the linguistic content and speaker\nidentity. In this paper, we study the disentanglement and recomposition of\nemotional elements in speech through variational autoencoding Wasserstein\ngenerative adversarial network (VAW-GAN). We propose a speaker-dependent EVC\nframework based on VAW-GAN, that includes two VAW-GAN pipelines, one for\nspectrum conversion, and another for prosody conversion. We train a spectral\nencoder that disentangles emotion and prosody (F0) information from spectral\nfeatures; we also train a prosodic encoder that disentangles emotion modulation\nof prosody (affective prosody) from linguistic prosody. At run-time, the\ndecoder of spectral VAW-GAN is conditioned on the output of prosodic VAW-GAN.\nThe vocoder takes the converted spectral and prosodic features to generate the\ntarget emotional speech. Experiments validate the effectiveness of our proposed\nmethod in both objective and subjective evaluations.\n