Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language Models

Humans can learn structural properties about a word from minimal experience,\nand deploy their learned syntactic representations uniformly in different\ngrammatical contexts. We assess the ability of modern neural language models to\nreproduce this behavior in English and evaluate the effect of structural\nsupervision on learning outcomes. First, we assess few-shot learning\ncapabilities by developing controlled experiments that probe models' syntactic\nnominal number and verbal argument structure generalizations for tokens seen as\nfew as two times during training. Second, we assess invariance properties of\nlearned representation: the ability of a model to transfer syntactic\ngeneralizations from a base context (e.g., a simple declarative active-voice\nsentence) to a transformed context (e.g., an interrogative sentence). We test\nfour models trained on the same dataset: an n-gram baseline, an LSTM, and two\nLSTM-variants trained with explicit structural supervision (Dyer et al.,2016;\nCharniak et al., 2016). We find that in most cases, the neural models are able\nto induce the proper syntactic generalizations after minimal exposure, often\nfrom just two examples during training, and that the two structurally\nsupervised models generalize more accurately than the LSTM model. All neural\nmodels are able to leverage information learned in base contexts to drive\nexpectations in transformed contexts, indicating that they have learned some\ninvariance properties of syntax.\n

Paper

References (29)

Scroll for more · 17 remaining

Similar papers

© 2026 NYSGPT2525 LLC