English
Related papers

Related papers: The JRC-Acquis: A multilingual aligned parallel co…

200 papers

EuroVoc (2012) is a highly multilingual thesaurus consisting of over 6,700 hierarchically organised subject domains used by European Institutions and many authorities in Member States of the European Union (EU) for the classification and…

Computation and Language · Computer Science 2013-09-23 Ralf Steinberger , Mohamed Ebrahim , Marco Turchi

We present the ELITR ECA corpus, a multilingual corpus derived from publications of the European Court of Auditors. We use automatic translation together with Bleualign to identify parallel sentence pairs in all 506 translation directions.…

Computation and Language · Computer Science 2021-09-16 Philip Williams , Barry Haddow

Cross-lingual annotations of legislative texts enable us to explore major themes covered in multilingual legal data and are a key facilitator of semantic similarity when searching for similar documents. Multilingual probabilistic topic…

Information Retrieval · Computer Science 2019-12-02 Carlos Badenes-Olmedo , Jose-Luis Redondo-Garcia , Oscar Corcho

Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at…

Computation and Language · Computer Science 2020-10-13 Ahmed El-Kishky , Vishrav Chaudhary , Francisco Guzman , Philipp Koehn

EuroVoc is a multilingual thesaurus that was built for organizing the legislative documentary of the European Union institutions. It contains thousands of categories at different levels of specificity and its descriptors are targeted by…

Computation and Language · Computer Science 2021-08-11 Andrei-Marius Avram , Vasile Pais , Dan Tufis

Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Although there are billions of parallel sentences for a few…

Computation and Language · Computer Science 2022-03-01 Makoto Morishita , Katsuki Chousa , Jun Suzuki , Masaaki Nagata

We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information…

Computation and Language · Computer Science 2025-09-09 Jinrui Yang , Timothy Baldwin , Trevor Cohn

Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no parallel text…

Computation and Language · Computer Science 2025-08-05 Motaz Saad , David Langlois , Kamel Smaili

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

Computation and Language · Computer Science 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

Existing summarization datasets come with two main drawbacks: (1) They tend to focus on overly exposed domains, such as news articles or wiki-like texts, and (2) are primarily monolingual, with few multilingual datasets. In this work, we…

Computation and Language · Computer Science 2022-10-25 Dennis Aumiller , Ashish Chouhan , Michael Gertz

With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological properties of languages is fundamental for progress in…

Computation and Language · Computer Science 2021-07-16 Ayyoob Imani , Masoud Jalili Sabet , Philipp Dufter , Michael Cysouw , Hinrich Schütze

Cross-lingual document classification aims at training a document classifier on resources in one language and transferring it to a different language without any additional resources. Several approaches have been proposed in the literature…

Computation and Language · Computer Science 2018-05-28 Holger Schwenk , Xian Li

The Scielo database is an important source of scientific information in Latin America, containing articles from several research domains. A striking characteristic of Scielo is that many of its full-text contents are presented in more than…

Computation and Language · Computer Science 2019-05-07 Felipe Soares , Viviane Pereira Moreira , Karin Becker

With the ongoing growth in number of digital articles in a wider set of languages and the expanding use of different languages, we need annotation methods that enable browsing multi-lingual corpora. Multilingual probabilistic topic models…

Computation and Language · Computer Science 2021-01-11 Carlos Badenes-Olmedo , Jose-Luis Redondo García , Oscar Corcho

We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several techniques to lower the…

Computation and Language · Computer Science 2020-07-08 Tom Kocmi , Martin Popel , Ondrej Bojar

This paper presents a language-independent approach to controlled vocabulary keyword assignment using the EUROVOC thesaurus. Due to the multilingual nature of EUROVOC, the keywords for a document written in one language can be displayed in…

Computation and Language · Computer Science 2007-05-23 Ralf Steinberger

We present a free Japanese-French parallel corpus. It includes 15M aligned segments and is obtained by compiling and filtering several existing resources. In this paper, we describe the existing resources, their quantity and quality, the…

Computation and Language · Computer Science 2022-08-30 Raoul Blin , Fabien Cromières

In this paper we describe the Japanese-English Subtitle Corpus (JESC). JESC is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue. It consists of more than 3.2 million examples, making…

Computation and Language · Computer Science 2018-02-22 Reid Pryzant , Yongjoo Chung , Dan Jurafsky , Denny Britz

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

Computation and Language · Computer Science 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer

We have developed an application aiming at federated search for EU and Hungarian legislation and jurisdiction. It now contains above 1 million documents, with daily updates. The database holds documents downloaded from the EU sources…

Computation and Language · Computer Science 2019-07-30 György Görög , Péter Weisz
‹ Prev 1 2 3 10 Next ›