Disfl-QA: A Benchmark Dataset for Understanding Disfluencies in Question Answering

Disfluencies is an under-studied topic in NLP, even though it is ubiquitous\nin human conversation. This is largely due to the lack of datasets containing\ndisfluencies. In this paper, we present a new challenge question answering\ndataset, Disfl-QA, a derivative of SQuAD, where humans introduce contextual\ndisfluencies in previously fluent questions. Disfl-QA contains a variety of\nchallenging disfluencies that require a more comprehensive understanding of the\ntext than what was necessary in prior datasets. Experiments show that the\nperformance of existing state-of-the-art question answering models degrades\nsignificantly when tested on Disfl-QA in a zero-shot setting.We show data\naugmentation methods partially recover the loss in performance and also\ndemonstrate the efficacy of using gold data for fine-tuning. We argue that we\nneed large-scale disfluency datasets in order for NLP models to be robust to\nthem. The dataset is publicly available at:\nhttps://github.com/google-research-datasets/disfl-qa.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC