End-to-end spoken language understanding using transformer networks and self-supervised pre-trained features

Transformer networks and self-supervised pre-training have consistently\ndelivered state-of-art results in the field of natural language processing\n(NLP); however, their merits in the field of spoken language understanding\n(SLU) still need further investigation. In this paper we introduce a modular\nEnd-to-End (E2E) SLU transformer network based architecture which allows the\nuse of self-supervised pre-trained acoustic features, pre-trained model\ninitialization and multi-task training. Several SLU experiments for predicting\nintent and entity labels/values using the ATIS dataset are performed. These\nexperiments investigate the interaction of pre-trained model initialization and\nmulti-task training with either traditional filterbank or self-supervised\npre-trained acoustic features. Results show not only that self-supervised\npre-trained acoustic features outperform filterbank features in almost all the\nexperiments, but also that when these features are used in combination with\nmulti-task training, they almost eliminate the necessity of pre-trained model\ninitialization.\n

Paper

References (27)

Scroll for more · 15 remaining

Similar papers

© 2026 NYSGPT2525 LLC