中文
相关论文

相关论文: Milimili. Collecting Parallel Data via Crowdsourci…

200 篇论文

Literature reviews allow scientists to stand on the shoulders of giants, showing promising directions, summarizing progress, and pointing out existing challenges in research. At the same time conducting a systematic literature review is a…

信息检索 · 计算机科学 2017-09-27 Evgeny Krivosheev , Fabio Casati , Valentina Caforio , Boualem Benatallah

Translation ambiguity, out of vocabulary words and missing some translations in bilingual dictionaries make dictionary-based Cross-language Information Retrieval (CLIR) a challenging task. Moreover, in agglutinative languages which do not…

信息检索 · 计算机科学 2014-11-06 Javid Dadashkarimi , Azadeh Shakery , Heshaam Faili

The task of organizing and clustering multilingual news articles for media monitoring is essential to follow news stories in real time. Most approaches to this task focus on high-resource languages (mostly English), with low-resource…

计算与语言 · 计算机科学 2022-04-29 João Santos , Afonso Mendes , Sebastião Miranda

The emergence of generative large language models (LLMs) raises the question: what will be its impact on crowdsourcing? Traditionally, crowdsourcing has been used for acquiring solutions to a wide variety of human-intelligence tasks,…

计算与语言 · 计算机科学 2023-10-23 Jan Cegin , Jakub Simko , Peter Brusilovsky

Large language models (LLMs) have demonstrated impressive translation capabilities even without being explicitly trained on parallel data. This remarkable property has led some to believe that parallel data is no longer necessary for…

计算与语言 · 计算机科学 2025-06-17 Muhammad Reza Qorib , Junyi Li , Hwee Tou Ng

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other…

We translate a closed text that is known in advance and available in many languages into a new and severely low resource language. Most human translation efforts adopt a portion-based approach to translate consecutive pages/chapters in…

计算与语言 · 计算机科学 2021-10-28 Zhong Zhou , Alex Waibel

Parallel corpora play an important role in training machine translation (MT) models, particularly for low-resource languages where high-quality bilingual data is scarce. This review provides a comprehensive overview of available parallel…

计算与语言 · 计算机科学 2025-04-23 Rahul Raja , Arpita Vats

In this paper we describe some ways to utilize various lexical resources to improve the quality of statistical machine translation system. We have augmented the training corpus with various lexical resources such as IndoWordnet semantic…

计算与语言 · 计算机科学 2017-03-07 Sreelekha S , Pushpak Bhattacharyya

Big data have the characteristics of enormous volume, high velocity, diversity, value-sparsity, and uncertainty, which lead the knowledge learning from them full of challenges. With the emergence of crowdsourcing, versatile information can…

机器学习 · 计算机科学 2022-06-22 Jing Zhang

Grammatical Error Correction (GEC) has been recently modeled using the sequence-to-sequence framework. However, unlike sequence transduction problems such as machine translation, GEC suffers from the lack of plentiful parallel data. We…

计算与语言 · 计算机科学 2019-04-12 Jared Lichtarge , Chris Alberti , Shankar Kumar , Noam Shazeer , Niki Parmar , Simon Tong

Simultaneous translation is vastly different from full-sentence translation, in the sense that it starts translation before the source sentence ends, with only a few words delay. However, due to the lack of large-scale, high-quality…

计算与语言 · 计算机科学 2021-09-24 Junkun Chen , Renjie Zheng , Atsuhito Kita , Mingbo Ma , Liang Huang

While crowdsourcing is an established solution for facilitating and scaling the collection of speech data, the involvement of non-experts necessitates protocols to ensure final data quality. To reduce the costs of these essential controls,…

计算与语言 · 计算机科学 2024-12-17 Beomseok Lee , Marco Gaido , Ioan Calapodescu , Laurent Besacier , Matteo Negri

We study a problem of optimal information gathering from multiple data providers that need to be incentivized to provide accurate information. This problem arises in many real world applications that rely on crowdsourced data sets, but…

计算机科学与博弈论 · 计算机科学 2017-11-27 Goran Radanovic , Adish Singla , Andreas Krause , Boi Faltings

Developing parallel corpora is an important and a difficult activity for Machine Translation. This requires manual annotation by Human Translators. Translating same text again is a useless activity. There are tools available to implement…

计算与语言 · 计算机科学 2012-10-23 Nisheeth Joshi , Iti Mathur

Interpreters facilitate multi-lingual meetings but the affordable set of languages is often smaller than what is needed. Automatic simultaneous speech translation can extend the set of provided languages. We investigate if such an automatic…

计算与语言 · 计算机科学 2021-06-18 Dominik Macháček , Matúš Žilinec , Ondřej Bojar

In this paper we present a novel method for retrieving information in languages other than that of the query. We use this technique in combination with existing traditional Cross Language Information Retrieval (CLIR) techniques to improve…

信息检索 · 计算机科学 2009-06-17 Mikhail Basilyan

Cross-Language Information Retrieval (CLIR) has become an important problem to solve in the recent years due to the growth of content in multiple languages in the Web. One of the standard methods is to use query translation from source to…

计算与语言 · 计算机科学 2016-08-05 Paheli Bhattacharya , Pawan Goyal , Sudeshna Sarkar

We present a novel method for obtaining high-quality, domain-targeted multiple choice questions from crowd workers. Generating these questions can be difficult without trading away originality, relevance or diversity in the answer options.…

人机交互 · 计算机科学 2017-07-20 Johannes Welbl , Nelson F. Liu , Matt Gardner

Cross-lingual summarization involves the summarization of text written in one language to a different one. There is a body of research addressing cross-lingual summarization from English to other European languages. In this work, we aim to…

计算与语言 · 计算机科学 2023-12-25 Nikhilesh Bhatnagar , Ashok Urlana , Vandan Mujadia , Pruthwik Mishra , Dipti Misra Sharma