This paper introduces voice reenactement as the task of voice conversion (VC)\nin which the expressivity of the source speaker is preserved during conversion\nwhile the identity of a target speaker is transferred. To do so, an original\nneural- VC architecture is proposed based on sequence-to-sequence voice\nconversion (S2S-VC) in which the speech prosody of the source speaker is\npreserved during conversion. First, the S2S-VC architecture is modified so as\nto synchronize the converted speech with the source speech by mean of phonetic\nduration encoding; second, the decoder is conditioned on the desired sequence\nof F0- values and an explicit F0-loss is formulated between the F0 of the\nsource speaker and the one of the converted speech. Besides, an adversarial\nlearning of conversions is integrated within the S2S-VC architecture so as to\nexploit both advantages of reconstruction of original speech and converted\nspeech with manipulated attributes during training and then reducing the\ninconsistency between training and conversion. An experimental evaluation on\nthe VCTK speech database shows that the speech prosody can be efficiently\npreserved during conversion, and that the proposed adversarial learning\nconsistently improves the conversion and the naturalness of the reenacted\nspeech.\n