Neural Module Networks (NMNs) aim at Visual Question Answering (VQA) via\ncomposition of modules that tackle a sub-task. NMNs are a promising strategy to\nachieve systematic generalization, i.e., overcoming biasing factors in the\ntraining distribution. However, the aspects of NMNs that facilitate systematic\ngeneralization are not fully understood. In this paper, we demonstrate that the\ndegree of modularity of the NMN have large influence on systematic\ngeneralization. In a series of experiments on three VQA datasets (VQA-MNIST,\nSQOOP, and CLEVR-CoGenT), our results reveal that tuning the degree of\nmodularity, especially at the image encoder stage, reaches substantially higher\nsystematic generalization. These findings lead to new NMN architectures that\noutperform previous ones in terms of systematic generalization.\n
Paper
References (27)
Scroll for more · 15 remaining