This paper presents a generic method for generating full facial 3D animation\nfrom speech. Existing approaches to audio-driven facial animation exhibit\nuncanny or static upper face animation, fail to produce accurate and plausible\nco-articulation or rely on person-specific models that limit their scalability.\nTo improve upon existing models, we propose a generic audio-driven facial\nanimation approach that achieves highly realistic motion synthesis results for\nthe entire face. At the core of our approach is a categorical latent space for\nfacial animation that disentangles audio-correlated and audio-uncorrelated\ninformation based on a novel cross-modality loss. Our approach ensures highly\naccurate lip motion, while also synthesizing plausible animation of the parts\nof the face that are uncorrelated to the audio signal, such as eye blinks and\neye brow motion. We demonstrate that our approach outperforms several baselines\nand obtains state-of-the-art quality both qualitatively and quantitatively. A\nperceptual user study demonstrates that our approach is deemed more realistic\nthan the current state-of-the-art in over 75% of cases. We recommend watching\nthe supplemental video before reading the paper:\nhttps://github.com/facebookresearch/meshtalk\n