中文
相关论文

相关论文: The ELITR ECA Corpus

200 篇论文

Canonical Correlation Analysis (CCA) has been exploited immensely for learning latent representations in various fields. This study takes a step further by demonstrating the potential of CCA in identifying Elementary Discourse Units(EDUs)…

计算与语言 · 计算机科学 2025-05-30 Akanksha Mehndiratta , Krishna Asawa

This paper introduces an updated and combined version of the bidirectional English-German EPIC-UdS (spoken) and EuroParl-UdS (written) corpora containing original European Parliament speeches as well as their translations and…

计算与语言 · 计算机科学 2026-03-17 Maria Kunilovskaya , Christina Pollkläsener

With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological properties of languages is fundamental for progress in…

计算与语言 · 计算机科学 2021-07-16 Ayyoob Imani , Masoud Jalili Sabet , Philipp Dufter , Michael Cysouw , Hinrich Schütze

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected…

The objective of this paper is to present a meta-corpus of diplomatic documents entitled Cartae Europae Medii Aevi or CEMA. It shows the logic and limits of this meta-corpus, which contains 250,000 documents, by specifying both its…

数字图书馆 · 计算机科学 2021-05-04 Nicolas Perreaux

We introduce KazParC, a parallel corpus designed for machine translation across Kazakh, English, Russian, and Turkish. The first and largest publicly available corpus of its kind, KazParC contains a collection of 371,902 parallel sentences…

计算与语言 · 计算机科学 2024-04-11 Rustem Yeshpanov , Alina Polonskaya , Huseyin Atakan Varol

The Parallel Meaning Bank is a corpus of translations annotated with shared, formal meaning representations comprising over 11 million words divided over four languages (English, German, Italian, and Dutch). Our approach is based on…

The European Commission's (EC) Directorate General for Translation, together with the EC's Joint Research Centre, is making available a large translation memory (TM; i.e. sentences and their professionally produced translations) covering…

计算与语言 · 计算机科学 2013-09-23 Ralf Steinberger , Andreas Eisele , Szymon Klocek , Spyridon Pilos , Patrick Schlüter

We introduce the Emergent Language Corpus Collection (ELCC): a collection of corpora generated from open source implementations of emergent communication systems across the literature. These systems include a variety of signalling game…

计算与语言 · 计算机科学 2024-12-05 Brendon Boldt , David Mortensen

In this paper, we present DIETA, a small, decoder-only Transformer model with 0.5 billion parameters, specifically designed and trained for Italian-English machine translation. We collect and curate a large parallel corpus consisting of…

计算与语言 · 计算机科学 2026-01-27 Pranav Kasela , Marco Braga , Alessandro Ghiotto , Andrea Pilzer , Marco Viviani , Alessandro Raganato

We present RONEC - the Named Entity Corpus for the Romanian language. The corpus contains over 26000 entities in ~5000 annotated sentences, belonging to 16 distinct classes. The sentences have been extracted from a copy-right free…

计算与语言 · 计算机科学 2020-04-28 Stefan Daniel Dumitrescu , Andrei-Marius Avram

We introduce S2ORC, a large corpus of 81.1M English-language academic papers spanning many academic disciplines. The corpus consists of rich metadata, paper abstracts, resolved bibliographic references, as well as structured full text for…

计算与语言 · 计算机科学 2020-07-08 Kyle Lo , Lucy Lu Wang , Mark Neumann , Rodney Kinney , Dan S. Weld

The patterns in which the syntax of different languages converges and diverges are often used to inform work on cross-lingual transfer. Nevertheless, little empirical work has been done on quantifying the prevalence of different syntactic…

计算与语言 · 计算机科学 2020-07-14 Dmitry Nikolaev , Ofir Arviv , Taelin Karidi , Neta Kenneth , Veronika Mitnik , Lilja Maria Saeboe , Omri Abend

Research on Large Language Models (LLMs) has recently witnessed an increasing interest in extending the models' context size to better capture dependencies within long documents. While benchmarks have been proposed to assess long-range…

计算与语言 · 计算机科学 2025-01-20 Thibaut Thonet , Jos Rozen , Laurent Besacier

We propose a method for efficiently finding all parallel passages in a large corpus, even if the passages are not quite identical due to rephrasing and orthographic variation. The key ideas are the representation of each word in the corpus…

计算与语言 · 计算机科学 2023-06-22 Avi Shmidman , Moshe Koppel , Ely Porat

Most judicial decisions involve the interpretation of legal texts; as such, judicial opinion requires the use of language as a medium to comment on or draw attention to other language. Language used this way is called metalanguage. We…

计算与语言 · 计算机科学 2023-05-25 Michael Kranzlein , Nathan Schneider , Kevin Tobia

Complete Multi-lingual Neural Machine Translation (C-MNMT) achieves superior performance against the conventional MNMT by constructing multi-way aligned corpus, i.e., aligning bilingual training examples from different language pairs when…

计算与语言 · 计算机科学 2024-07-23 Yulin Xu , Zhen Yang , Fandong Meng , JieZhou

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese research community.…

计算与语言 · 计算机科学 2022-10-20 Chinh Ngo , Trieu H. Trinh , Long Phan , Hieu Tran , Tai Dang , Hieu Nguyen , Minh Nguyen , Minh-Thang Luong

We introduce MULTI-EURLEX, a new multilingual dataset for topic classification of legal documents. The dataset comprises 65k European Union (EU) laws, officially translated in 23 languages, annotated with multiple labels from the EUROVOC…

计算与语言 · 计算机科学 2021-09-08 Ilias Chalkidis , Manos Fergadiotis , Ion Androutsopoulos

This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of speech) which have…