How Multilingual Are Large Language Models Fine-Tuned for Translation?

A new paradigm for machine translation has recently emerged: fine-tuning large language models (LLM) on parallel text has been shown to outperform dedicated translation systems trained in a supervised fashion on much larger amounts of parallel data (Xu et al., 2024a; Alves et al., 2024). However, it remains unclear whether this paradigm can enable massively multilingual machine translation or whether it requires fine-tuning dedicated models for a small number of language pairs. How does translation fine-tuning impact the MT capabilities of LLMs for zero-shot languages, zero-shot language pairs, and translation tasks that do not involve English? To address these questions, we conduct an extensive empirical evaluation of the translation quality of the TOWER family of language models (Alves et al., 2024) on 132 translation tasks from the multi-parallel FLORES-200 data. We find that translation fine-tuning improves translation quality even for zero-shot languages on average, but that the impact is uneven depending on the language pairs involved. These results call for further research to effectively enable massively multilingual translation with LLMs.

Paper

References (37)

Scroll for more · 25 remaining

Similar papers

Reviewer sAju7/10 · confidence 4/52024-05-02

Summary

This paper investigates the capabilities of a large language model (LLM) fine-tuned for machine translation (MT) on certain languages to perform across various language directions for which it has not been tuned. This paper does not present any technical novelty but tries to understand better the capabilities of an LLM challenging it in unseen language directions. The experiments include 132 language directions split into fully supervised, partially supervised, and unsupervised directions, depending on the data used during training. When dealing with a large number of directions, how to report the results is an important problem, because the widely-used average does not allow a reader to highlight the positive and negative performance of a model. This paper addresses this problem by considering the best and worst behaviors of a model and linking them with the source and target languages. A large set of experiments with in-depth analysis are reported to verify the initial research questions. The main findings of the paper are: -) The fine-tuning of an LLM for translation-related tasks helps to improve the performance also on unseen languages. -) The paper identifies high variance in translation quality for language pairs that are not supervised and outliers languages such as Icelandic and Korean.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

Although it does not propose any technical enhancement, the paper addresses an interesting problem that should help to understand better the behaviors of LLMs for translation. In this historical period, where there is an increasing emphasis on improving LLM performance on multiple tasks, these papers are important to have a critical view of the existing approaches and provide some directions to work on. The paper is supported by a large set of experiments that include results for high- and low-resourced languages. The proposed analysis has clear questions that are addressed with proper experiments. The paper is clear and easy to follow.

Reasons to reject

The findings are clear but not conclusive leaving more activities to better understand the capabilities of the LLMs for translation and to motivate the outlier behaviors of some languages. The proposed method for overcoming the average limitation when reporting results for a large number of pairs is interesting, but it may not be immediately comprehensible and may require some effort to fully appreciate. However, this is the trade-off for introducing a new metric or visualization strategy. The experimental parts of the paper are not always easy to follow and appreciate. The fact that the figures and tables are not on the page where they are discussed creates more difficulties. Overall the paper is well-written, but it lacks the final proofreading. Some sentences to adjust: -) Page 1: “Briakou et al. (2023). vast amounts of noisy multilingual pre-training data. Xu” not clear how the sentence starting with “vast” is connected to the previous and next sentences. -) Page 3: the table with the language information is table 1 and not table 2. -) Page 5: “for each target language for the for the dedicated” “For the” is repeated. -) Page 8: “upsample lower resource languages .“ extra space to be removed.

Questions to authors

Do the authors have any insights on the impact of the alphabet used by each language on the translation performance? And what is the role of the vocabulary size (smaller vocabulary can penalize those languages with a high-variable alphabet or an alphabet with a large set of specific symbols)?

Reviewer zRXH6/10 · confidence 3/52024-05-11

Summary

The paper conducts an extensive empirical evaluation of the translation quality of the TOWER family of language models on 132 translation tasks from the multi-parallel FLORES data, aiming to find out how does translation fine-tuning impact the MT capabilities of LLMs for zero-shot languages, zero-shot language pairs, and translation tasks that do not involve English. The conclusion is interesting that the translation fine-tuning improves translation quality even for zero-shot languages on average.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1. The paper addresses an interesting problem about the multilingual translation ability of LLM models. 2. Empirical experiments show that the translation fine-tuning improves translation quality even for zero-shot languages on average, which is impressive. 3. The paper further finds out that the improves are not defined by the language similarity, which leaves rooms for the future research on this direction.

Reasons to reject

1. The paper seems to be written in hurry. Table 2 (ref in Section 2) seems to be missing. Figure 5 is cropped. 2. The innovation of the paper seems to be limited. However, it is ok for an empirical paper.

Questions to authors

It will be better if the finding can be further confirmed on other LLMs.

Reviewer rg127/10 · confidence 4/52024-05-12

Summary

The authors conduct a large scale empirical evaluation of the translation performance of the Tower language models, which are multilingual foundation models created from the LLama2 line of base models, and targeted at translation and related tasks such as quality estimation and automatic post editing. The evaluation measures translation performance in supervised and zero-shot settings, before and after fine-tuning. As a baseline and semi-oracle, the NLLB massively multilingual model is also evaluated in all settings.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

- clearly written paper with very useful evaluation on partial- and fully-unsupervised MT tasks. - in particular, the inclusion of best and worst performance in each setting will be useful for many researchers - concrete takeaways, such as that pretraining on translation tasks benefits translation performance on unseen language pairs, but not for all language pairs uniformly/

Reasons to reject

- limited scientific novelty, in that the work is only a collection of evaluations

Reviewer sAju2024-06-03

Acknowledge

Thanks a lot for the clarification

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC