English
Related papers

Related papers: Open Subtitles Paraphrase Corpus for Six Languages

200 papers

This paper presents a high-quality multilingual dataset for the documentation domain to advance research on localization of structured text. Unlike widely-used datasets for translation of plain text, we collect XML-structured parallel text…

Computation and Language · Computer Science 2020-06-25 Kazuma Hashimoto , Raffaella Buschiazzo , James Bradbury , Teresa Marshall , Richard Socher , Caiming Xiong

ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all together 6 thousand hours in size. The corpora were built in an automatic fashion from the…

Computation and Language · Computer Science 2026-04-16 Nikola Ljubešić , Peter Rupnik , Ivan Porupski , Taja Kuzman Pungeršek

This paper accompanies the software documentation data set for machine translation, a parallel evaluation data set of data originating from the SAP Help Portal, that we released to the machine translation community for research purposes. It…

Computation and Language · Computer Science 2020-11-13 Bianka Buschbeck , Miriam Exel

Research in question answering datasets and models has gained a lot of attention in the research community. Many of them release their own question answering datasets as well as the models. There is tremendous progress that we have seen in…

Computation and Language · Computer Science 2021-12-28 Andreas Chandra , Affandy Fahrizain , Ibrahim , Simon Willyanto Laufried

The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME (Biblioteca Regional de Medicina) in agreement with the Pan American Health…

Computation and Language · Computer Science 2019-05-07 Felipe Soares , Martin Krallinger

Existing paraphrase identification datasets lack sentence pairs that have high lexical overlap without being paraphrases. Models trained on such data fail to distinguish pairs like flights from New York to Florida and flights from Florida…

Computation and Language · Computer Science 2019-04-03 Yuan Zhang , Jason Baldridge , Luheng He

Manual correction of speech transcription can involve a selection from plausible transcriptions. Recent work has shown the feasibility of employing a mismatched crowd for speech transcription. However, it is yet to be established whether a…

Artificial Intelligence · Computer Science 2016-09-08 Purushotam Radadia , Shirish Karande

When translating phrases (words or group of words), human translators, consciously or not, resort to different translation processes apart from the literal translation, such as Idiom Equivalence, Generalization, Particularization, Semantic…

Computation and Language · Computer Science 2019-04-30 Yuming Zhai , Pooyan Safari , Gabriel Illouz , Alexandre Allauzen , Anne Vilnat

This paper examines the current state-of-the-art of German text simplification, focusing on parallel and monolingual German corpora. It reviews neural language models for simplifying German texts and assesses their suitability for legal…

Computation and Language · Computer Science 2023-12-18 Thorben Schomacker , Michael Gille , Jörg von der Hülls , Marina Tropmann-Frick

Machine translation systems achieve near human-level performance on some languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentences, which hinders their applicability to the majority of…

Computation and Language · Computer Science 2018-08-15 Guillaume Lample , Myle Ott , Alexis Conneau , Ludovic Denoyer , Marc'Aurelio Ranzato

We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United States from 2000 to 2021.…

Computation and Language · Computer Science 2025-08-25 Masaaki Nagata , Katsuki Chousa , Norihito Yasuda

We learn a joint multilingual sentence embedding and use the distance between sentences in different languages to filter noisy parallel data and to mine for parallel data in large news collections. We are able to improve a competitive…

Computation and Language · Computer Science 2018-05-28 Holger Schwenk

Recent progress in semantic parsing scarcely considers languages other than English but professional translation can be prohibitively expensive. We adapt a semantic parser trained on a single language, such as English, to new languages and…

Computation and Language · Computer Science 2020-09-24 Tom Sherborne , Yumo Xu , Mirella Lapata

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

Computation and Language · Computer Science 2023-08-25 Emily Silcock , Melissa Dell

The number of senses of a given word, or polysemy, is a very subjective notion, which varies widely across annotators and resources. We propose a novel method to estimate polysemy, based on simple geometry in the contextual embedding space.…

Computation and Language · Computer Science 2023-05-03 Christos Xypolopoulos , Antoine J. -P. Tixier , Michalis Vazirgiannis

Spoken Language Understanding (SLU) is one of the core components of a task-oriented dialogue system, which aims to extract the semantic meaning of user queries (e.g., intents and slots). In this work, we introduce OpenSLU, an open-source…

Computation and Language · Computer Science 2023-05-18 Libo Qin , Qiguang Chen , Xiao Xu , Yunlong Feng , Wanxiang Che

Being able to understand information is a key factor for a self-determined life and society. It is also very important for participating in democratic processes. The study of automatic text simplification is often limited by the…

Computation and Language · Computer Science 2026-03-17 Stefan Bott , Verena Riegler , Horacio Saggion , Almudena Rascón Alcaina , Nouran Khallaf

Despite impressive advancements in multilingual corpora collection and model training, developing large-scale deployments of multilingual models still presents a significant challenge. This is particularly true for language tasks that are…

Computation and Language · Computer Science 2023-06-14 Łukasz Augustyniak , Szymon Woźniak , Marcin Gruza , Piotr Gramacki , Krzysztof Rajda , Mikołaj Morzy , Tomasz Kajdanowicz

Translation systems, including foundation models capable of translation, can produce errors that result in gender mistranslation, and such errors can be especially harmful. To measure the extent of such potential harms when translating into…

Computation and Language · Computer Science 2024-10-07 Kevin Robinson , Sneha Kudugunta , Romina Stella , Sunipa Dev , Jasmijn Bastings

We describe efforts towards getting better resources for English-Arabic machine translation of spoken text. In particular, we look at movie subtitles as a unique, rich resource, as subtitles in one language often get translated into other…

Computation and Language · Computer Science 2016-09-06 Fahad Al-Obaidli , Stephen Cox , Preslav Nakov