逆势而上:面向统计机器翻译的大规模成本聚焦主动学习
计算与语言
2014-10-23 v1 机器学习
机器学习
摘要
我们探索在已拥有大量资源的情况下,如何通过添加更多翻译数据来改进机器翻译系统。主要挑战在于如何扭转通常遇到的收益递减趋势。我们提出了一种主动学习式的数据征集算法来应对这一挑战。我们通过 Amazon Mechanical Turk 收集标注进行测试,发现性能提升速率获得了一个数量级的增长。
引用
@article{arxiv.1410.5877,
title = {Bucking the Trend: Large-Scale Cost-Focused Active Learning for Statistical Machine Translation},
author = {Michael Bloodgood and Chris Callison-Burch},
journal= {arXiv preprint arXiv:1410.5877},
year = {2014}
}
备注
11 pages, 14 figures; appeared in Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, July 2010