Pretrained Semantic Speech Embeddings for End-to-End Spoken Language Understanding via Cross-Modal Teacher-Student Learning
Spoken language understanding is typically based on pipeline architectures\nincluding speech recognition and natural language understanding steps. These\ncomponents are optimized independently to allow usage of available data, but\nthe overall system suffers from error propagation. In this paper, we propose a\nnovel training method that enables pretrained contextual embeddings to process\nacoustic features. In particular, we extend it with an encoder of pretrained\nspeech recognition systems in order to construct end-to-end spoken language\nunderstanding systems. Our proposed method is based on the teacher-student\nframework across speech and text modalities that aligns the acoustic and the\nsemantic latent spaces. Experimental results in three benchmarks show that our\nsystem reaches the performance comparable to the pipeline architecture without\nusing any training data and outperforms it after fine-tuning with ten examples\nper class on two out of three benchmarks.\n