Improving Compositional Generalization with Latent Structure and Data Augmentation

Generic unstructured neural networks have been shown to struggle on\nout-of-distribution compositional generalization. Compositional data\naugmentation via example recombination has transferred some prior knowledge\nabout compositionality to such black-box neural models for several semantic\nparsing tasks, but this often required task-specific engineering or provided\nlimited gains.\n We present a more powerful data recombination method using a model called\nCompositional Structure Learner (CSL). CSL is a generative model with a\nquasi-synchronous context-free grammar backbone, which we induce from the\ntraining data. We sample recombined examples from CSL and add them to the\nfine-tuning data of a pre-trained sequence-to-sequence model (T5). This\nprocedure effectively transfers most of CSL's compositional bias to T5 for\ndiagnostic tasks, and results in a model even stronger than a T5-CSL ensemble\non two real world compositional generalization tasks. This results in new\nstate-of-the-art performance for these challenging semantic parsing tasks\nrequiring generalization to both natural language variation and novel\ncompositions of elements.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC