Rethinking End-to-End Evaluation of Decomposable Tasks: A Case Study on Spoken Language Understanding
Decomposable tasks are complex and comprise of a hierarchy of sub-tasks.\nSpoken intent prediction, for example, combines automatic speech recognition\nand natural language understanding. Existing benchmarks, however, typically\nhold out examples for only the surface-level sub-task. As a result, models with\nsimilar performance on these benchmarks may have unobserved performance\ndifferences on the other sub-tasks. To allow insightful comparisons between\ncompetitive end-to-end architectures, we propose a framework to construct\nrobust test sets using coordinate ascent over sub-task specific utility\nfunctions. Given a dataset for a decomposable task, our method optimally\ncreates a test set for each sub-task to individually assess sub-components of\nthe end-to-end model. Using spoken language understanding as a case study, we\ngenerate new splits for the Fluent Speech Commands and Snips SmartLights\ndatasets. Each split has two test sets: one with held-out utterances assessing\nnatural language understanding abilities, and one with held-out speakers to\ntest speech processing skills. Our splits identify performance gaps up to 10%\nbetween end-to-end systems that were within 1% of each other on the original\ntest sets. These performance gaps allow more realistic and actionable\ncomparisons between different architectures, driving future model development.\nWe release our splits and tools for the community.\n
Paper
References (36)
Scroll for more · 24 remaining