Modality Dropout for Improved Performance-driven Talking Faces

We describe our novel deep learning approach for driving animated faces using both acoustic and visual information. In particular, speech-related facial movements are generated using audiovisual information, and non-speech facial movements are generated using only visual information. To ensure that our model exploits both modalities during training, batches are generated that contain audio-only…

Paper

Similar papers

© 2026 NYSGPT2525 LLC