Efficient data selection for machine translation

2008 IEEE Spoken Language Technology Workshop Pub Date : 2008-12-01 DOI:10.1109/SLT.2008.4777890

Arindam Mandal, D. Vergyri, Wen Wang, Jing Zheng, A. Stolcke, Gökhan Tür, Dilek Z. Hakkani-Tür, N. F. Ayan

引用次数: 27

Abstract

Performance of statistical machine translation (SMT) systems relies on the availability of a large parallel corpus which is used to estimate translation probabilities. However, the generation of such corpus is a long and expensive process. In this paper, we introduce two methods for efficient selection of training data to be translated by humans. Our methods are motivated by active learning and aim to choose new data that adds maximal information to the currently available data pool. The first method uses a measure of disagreement between multiple SMT systems, whereas the second uses a perplexity criterion. We performed experiments on Chinese-English data in multiple domains and test sets. Our results show that we can select only one-fifth of the additional training data and achieve similar or better translation performance, compared to that of using all available data.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

高效的机器翻译数据选择

统计机器翻译(SMT)系统的性能依赖于用于估计翻译概率的大型并行语料库的可用性。然而，这种语料库的生成是一个漫长而昂贵的过程。本文介绍了两种人工翻译训练数据的有效选择方法。我们的方法以主动学习为动力，旨在选择向当前可用数据池中添加最大信息的新数据。第一种方法使用多个SMT系统之间的分歧度量，而第二种方法使用困惑标准。我们在多个领域和测试集上对汉英数据进行了实验。我们的结果表明，与使用所有可用数据相比，我们可以只选择五分之一的额外训练数据并获得类似或更好的翻译性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2008 IEEE Spoken Language Technology Workshop

自引率

0.00%

发文量

期刊最新文献

“Who is this” quiz dialogue system and users' evaluation Latent dirichlet language model for speech recognition Modelling user behaviour in the HIS-POMDP dialogue manager A syntactic language model based on incremental CCG parsing Improving word segmentation for Thai speech translation