How does the level of human input affect the quality of automated legal translation? A case study comparing four workflows of human-AI collaboration
This study explores the impact of human-AI collaboration on the quality of AI-translated legal documents. Focusing on Chinese–English legal translations produced by a neural machine translation (NMT) system and a Large Language Model (LLM), the study compares four translation workflows with varying degrees of human-AI collaboration: (1) no human input: NMT translation alone; (2) low human input: LLM translation with minimal human instruction; (3) medium human input: LLM translation with context-specific instructions; and (4) high human input: LLM translation with context-specific instructions and a human-verified glossary. The translations were evaluated using automatic metrics (BLEU and chrF++) and manual error analysis. Results show that automated legal translation quality improves with increased levels of human input, with the best quality achieved from a detailed prompt carrying context-specific instructions and a bilingual glossary. However, despite the quality improvements due to increased human input, all the automated translations still fell short of professional human translations, especially in handling specialized legal terminology and complex syntax. These results suggest that full automation is not yet feasible in high-stakes domains such as legal translation, underscoring the irreplaceability of human translators and the importance of human-AI collaboration in guaranteeing adequate translation quality in the legal context.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex