Fine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic Modeling
Neural topic models can augment or replace bag-of-words inputs with the\nlearned representations of deep pre-trained transformer-based word prediction\nmodels. One added benefit when using representations from multilingual models\nis that they facilitate zero-shot polylingual topic modeling. However, while it\nhas been widely observed that pre-trained embeddings should be fine-tuned to a\ngiven task, it is not immediately clear what supervision should look like for\nan unsupervised task such as topic modeling. Thus, we propose several methods\nfor fine-tuning encoders to improve both monolingual and zero-shot polylingual\nneural topic modeling. We consider fine-tuning on auxiliary tasks, constructing\na new topic classification task, integrating the topic classification objective\ndirectly into topic model training, and continued pre-training. We find that\nfine-tuning encoder representations on topic classification and integrating the\ntopic classification task directly into topic modeling improves topic quality,\nand that fine-tuning encoder representations on any task is the most important\nfactor for facilitating cross-lingual transfer.\n