Choose Your QA Model Wisely: A Systematic Study of Generative and Extractive Readers for Question Answering

While both extractive and generative readers have been successfully applied\nto the Question Answering (QA) task, little attention has been paid toward the\nsystematic comparison of them. Characterizing the strengths and weaknesses of\nthe two readers is crucial not only for making a more informed reader selection\nin practice but also for developing a deeper understanding to foster further\nresearch on improving readers in a principled manner. Motivated by this goal,\nwe make the first attempt to systematically study the comparison of extractive\nand generative readers for question answering. To be aligned with the\nstate-of-the-art, we explore nine transformer-based large pre-trained language\nmodels (PrLMs) as backbone architectures. Furthermore, we organize our findings\nunder two main categories: (1) keeping the architecture invariant, and (2)\nvarying the underlying PrLMs. Among several interesting findings, it is\nimportant to highlight that (1) the generative readers perform better in long\ncontext QA, (2) the extractive readers perform better in short context while\nalso showing better out-of-domain generalization, and (3) the encoder of\nencoder-decoder PrLMs (e.g., T5) turns out to be a strong extractive reader and\noutperforms the standard choice of encoder-only PrLMs (e.g., RoBERTa). We also\nstudy the effect of multi-task learning on the two types of readers varying the\nunderlying PrLMs and perform qualitative and quantitative diagnosis to provide\nfurther insights into future directions in modeling better readers.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC