Transcoding compositionally: using attention to find more generalizable solutions

While sequence-to-sequence models have shown remarkable generalization power\nacross several natural language tasks, their construct of solutions are argued\nto be less compositional than human-like generalization. In this paper, we\npresent seq2attn, a new architecture that is specifically designed to exploit\nattention to find compositional patterns in the input. In seq2attn, the two\nstandard components of an encoder-decoder model are connected via a transcoder,\nthat modulates the information flow between them. We show that seq2attn can\nsuccessfully generalize, without requiring any additional supervision, on two\ntasks which are specifically constructed to challenge the compositional skills\nof neural networks. The solutions found by the model are highly interpretable,\nallowing easy analysis of both the types of solutions that are found and\npotential causes for mistakes. We exploit this opportunity to introduce a new\nparadigm to test compositionality that studies the extent to which a model\novergeneralizes when confronted with exceptions. We show that seq2attn exhibits\nsuch overgeneralization to a larger degree than a standard sequence-to-sequence\nmodel.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC