Parameter Space Factorization for Zero-Shot Learning across Tasks and Languages

Most combinations of NLP tasks and language varieties lack in-domain examples\nfor supervised training because of the paucity of annotated data. How can\nneural models make sample-efficient generalizations from task-language\ncombinations with available data to low-resource ones? In this work, we propose\na Bayesian generative model for the space of neural parameters. We assume that\nthis space can be factorized into latent variables for each language and each\ntask. We infer the posteriors over such latent variables based on data from\nseen task-language combinations through variational inference. This enables\nzero-shot classification on unseen combinations at prediction time. For\ninstance, given training data for named entity recognition (NER) in Vietnamese\nand for part-of-speech (POS) tagging in Wolof, our model can perform accurate\npredictions for NER in Wolof. In particular, we experiment with a typologically\ndiverse sample of 33 languages from 4 continents and 11 families, and show that\nour model yields comparable or better results than state-of-the-art, zero-shot\ncross-lingual transfer methods. Moreover, we demonstrate that approximate\nBayesian model averaging results in smoother predictive distributions, whose\nentropy inversely correlates with accuracy. Hence, the proposed framework also\noffers robust estimates of prediction uncertainty. Our code is located at\ngithub.com/cambridgeltl/parameter-factorization\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC