Encodings of Source Syntax: Similarities in NMT Representations Across Target Languages

We train neural machine translation (NMT) models from English to six target\nlanguages, using NMT encoder representations to predict ancestor constituent\nlabels of source language words. We find that NMT encoders learn similar source\nsyntax regardless of NMT target language, relying on explicit morphosyntactic\ncues to extract syntactic features from source sentences. Furthermore, the NMT\nencoders outperform RNNs trained directly on several of the constituent label\nprediction tasks, suggesting that NMT encoder representations can be used\neffectively for natural language tasks involving syntax. However, both the NMT\nencoders and the directly-trained RNNs learn substantially different syntactic\ninformation from a probabilistic context-free grammar (PCFG) parser. Despite\nlower overall accuracy scores, the PCFG often performs well on sentences for\nwhich the RNN-based models perform poorly, suggesting that RNN architectures\nare constrained in the types of syntax they can learn.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC