We propose Deep Autoencoding Predictive Components (DAPC) -- a\nself-supervised representation learning method for sequence data, based on the\nintuition that useful representations of sequence data should exhibit a simple\nstructure in the latent space. We encourage this latent structure by maximizing\nan estimate of predictive information of latent feature sequences, which is the\nmutual information between past and future windows at each time step. In\ncontrast to the mutual information lower bound commonly used by contrastive\nlearning, the estimate of predictive information we adopt is exact under a\nGaussian assumption. Additionally, it can be computed without negative\nsampling. To reduce the degeneracy of the latent space extracted by powerful\nencoders and keep useful information from the inputs, we regularize predictive\ninformation learning with a challenging masked reconstruction loss. We\ndemonstrate that our method recovers the latent space of noisy dynamical\nsystems, extracts predictive features for forecasting tasks, and improves\nautomatic speech recognition when used to pretrain the encoder on large amounts\nof unlabeled data.\n
Paper
References (64)
Scroll for more · 38 remaining