Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders
Discrete latent spaces in variational autoencoders have been shown to\neffectively capture the data distribution for many real-world problems such as\nnatural language understanding, human intent prediction, and visual scene\nrepresentation. However, discrete latent spaces need to be sufficiently large\nto capture the complexities of real-world data, rendering downstream tasks\ncomputationally challenging. For instance, performing motion planning in a\nhigh-dimensional latent representation of the environment could be intractable.\nWe consider the problem of sparsifying the discrete latent space of a trained\nconditional variational autoencoder, while preserving its learned\nmultimodality. As a post hoc latent space reduction technique, we use\nevidential theory to identify the latent classes that receive direct evidence\nfrom a particular input condition and filter out those that do not. Experiments\non diverse tasks, such as image generation and human behavior prediction,\ndemonstrate the effectiveness of our proposed technique at reducing the\ndiscrete latent sample space size of a model while maintaining its learned\nmultimodality.\n