Does syntax need to grow on trees? Sources of hierarchical inductive bias in sequence-to-sequence networks

Learners that are exposed to the same training data might generalize\ndifferently due to differing inductive biases. In neural network models,\ninductive biases could in theory arise from any aspect of the model\narchitecture. We investigate which architectural factors affect the\ngeneralization behavior of neural sequence-to-sequence models trained on two\nsyntactic tasks, English question formation and English tense reinflection. For\nboth tasks, the training set is consistent with a generalization based on\nhierarchical structure and a generalization based on linear order. All\narchitectural factors that we investigated qualitatively affected how models\ngeneralized, including factors with no clear connection to hierarchical\nstructure. For example, LSTMs and GRUs displayed qualitatively different\ninductive biases. However, the only factor that consistently contributed a\nhierarchical bias across tasks was the use of a tree-structured model rather\nthan a model with sequential recurrence, suggesting that human-like syntactic\ngeneralization requires architectural syntactic structure.\n

Paper

References (54)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC