Machine translation has evolved rapidly along with the development of statistics and linguistics. The acquisition of machine translation knowledge is based on the existing language, corpus and machine learning methods, from which knowledge is extracted to improve the translation effect. The aim of this article is to study artificial intelligence machine automatic translation systems based on parallel corpora. The article describes the process of building a multilingual-based parallel corpus and gives a CAT software package. The weighting of this feature is obtained by defining the evaluation features of multiple parallel corpora and combining each evaluation feature using a linear model using the Perceptron algorithm. By analyzing the coverage contribution factors and quality assessment features of the packets, a subset of data suitable for statistical machine translation is proposed, and the decoding cost of the system is reduced to some extent.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex