From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding

The lack of publicly available evaluation data for low-resource languages\nlimits progress in Spoken Language Understanding (SLU). As key tasks like\nintent classification and slot filling require abundant training data, it is\ndesirable to reuse existing data in high-resource languages to develop models\nfor low-resource scenarios. We introduce xSID, a new benchmark for\ncross-lingual Slot and Intent Detection in 13 languages from 6 language\nfamilies, including a very low-resource dialect. To tackle the challenge, we\npropose a joint learning approach, with English SLU training data and\nnon-English auxiliary tasks from raw text, syntax and translation for transfer.\nWe study two setups which differ by type and language coverage of the\npre-trained embeddings. Our results show that jointly learning the main tasks\nwith masked language modeling is effective for slots, while machine translation\ntransfer works best for intent classification.\n

Paper

References (56)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC