SLUE: New Benchmark Tasks for Spoken Language Understanding Evaluation on Natural Speech

Progress in speech processing has been facilitated by shared datasets and\nbenchmarks. Historically these have focused on automatic speech recognition\n(ASR), speaker identification, or other lower-level tasks. Interest has been\ngrowing in higher-level spoken language understanding tasks, including using\nend-to-end models, but there are fewer annotated datasets for such tasks. At\nthe same time, recent work shows the possibility of pre-training generic\nrepresentations and then fine-tuning for several tasks using relatively little\nlabeled data. We propose to create a suite of benchmark tasks for Spoken\nLanguage Understanding Evaluation (SLUE) consisting of limited-size labeled\ntraining sets and corresponding evaluation sets. This resource would allow the\nresearch community to track progress, evaluate pre-trained representations for\nhigher-level tasks, and study open questions such as the utility of pipeline\nversus end-to-end approaches. We present the first phase of the SLUE benchmark\nsuite, consisting of named entity recognition, sentiment analysis, and ASR on\nthe corresponding datasets. We focus on naturally produced (not read or\nsynthesized) speech, and freely available datasets. We provide new\ntranscriptions and annotations on subsets of the VoxCeleb and VoxPopuli\ndatasets, evaluation metrics and results for baseline models, and an\nopen-source toolkit to reproduce the baselines and evaluate new models.\n

Paper

References (45)

Scroll for more · 33 remaining

Similar papers

© 2026 NYSGPT2525 LLC