The NiuTrans Machine Translation Systems for WMT22
Published in Workshop, 2022
This paper describes the NiuTrans neural machine translation systems of the WMT22 General MT constrained task. We participate in four directions, including Chinese→English, English→Croatian, and Livonian↔English. Our models are based on several advanced Transformer variants, e.g., Transformer-ODE, Universal Multiscale Transformer (UMST). The main workflow consists of data filtering, large-scale data augmentation (i.e., iterative back-translation, iterative knowledge distillation), and specific-domain fine-tuning. Moreover, we try several multi-domain methods, such as a multi-domain model structure and a multi-domain data clustering method, to rise to this year’s newly proposed multi-domain test set challenge. For low-resource scenarios, we build a multi-language translation model to enhance the performance and try to use the pretrained language model (mBERT) to initialize the translation model.
Recommended citation: Zhang Y, Wang Z, Cao R, et al. The niutrans machine translation systems for wmt20[C]//Proceedings of the Fifth Conference on Machine Translation. 2020: 338-345.
Download Paper
