Dyna-bAbI: unlocking bAbI's potential with dynamic synthetic benchmarking

While neural language models often perform surprisingly well on natural\nlanguage understanding (NLU) tasks, their strengths and limitations remain\npoorly understood. Controlled synthetic tasks are thus an increasingly\nimportant resource for diagnosing model behavior. In this work we focus on\nstory understanding, a core competency for NLU systems. However, the main\nsynthetic resource for story understanding, the bAbI benchmark, lacks such a\nsystematic mechanism for controllable task generation. We develop Dyna-bAbI, a\ndynamic framework providing fine-grained control over task generation in bAbI.\nWe demonstrate our ideas by constructing three new tasks requiring\ncompositional generalization, an important evaluation setting absent from the\noriginal benchmark. We tested both special-purpose models developed for bAbI as\nwell as state-of-the-art pre-trained methods, and found that while both\napproaches solve the original tasks (>99% accuracy), neither approach succeeded\nin the compositional generalization setting, indicating the limitations of the\noriginal training data. We explored ways to augment the original data, and\nfound that though diversifying training data was far more useful than simply\nincreasing dataset size, it was still insufficient for driving robust\ncompositional generalization (with <70% accuracy for complex compositions). Our\nresults underscore the importance of highly controllable task generators for\ncreating robust NLU systems through a virtuous cycle of model and data\ndevelopment.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC