English
Related papers

Related papers: HELFI: a Hebrew-Greek-Finnish Parallel Bible Corpu…

200 papers

The development of bilingual dictionaries to medieval translations presents diverse difficulties. These result from two types of philological circumstances: a) the asymmetry between the source language and the target language; and b) the…

Computation and Language · Computer Science 2023-08-29 Martin Ruskov , Lora Taseva

The Linguistic Annotation Framework (LAF) provides a general, extensible stand-off markup system for corpora. This paper discusses LAF-Fabric, a new tool to analyse LAF resources in general with an extension to process the Hebrew Bible in…

Computation and Language · Computer Science 2015-01-13 Dirk Roorda , Gino Kalkman , Martijn Naaijer , Andreas van Cranenburgh

The systematic study of ancient texts including their production, transmission and interpretation is greatly aided by the digital methods that started taking off in the 1970s. But how is that research in turn transmitted to new generations…

Computation and Language · Computer Science 2015-01-09 Dirk Roorda

In the prose style transfer task a system, provided with text input and a target prose style, produces output which preserves the meaning of the input text but alters the style. These systems require parallel data for evaluation of results…

Computation and Language · Computer Science 2021-09-01 Keith Carlson , Allen Riddell , Daniel Rockmore

This contribution to a special issue on "Computer-aided processing of intertextuality" in ancient texts will illustrate how using digital tools to interact with the Hebrew Bible offers new promising perspectives for visualizing the texts…

Computation and Language · Computer Science 2017-10-25 Nicolai Winther-Nielsen

One of the primary tasks of morphological parsers is the disambiguation of homographs. Particularly difficult are cases of unbalanced ambiguity, where one of the possible analyses is far more frequent than the others. In such cases, there…

Computation and Language · Computer Science 2020-10-07 Avi Shmidman , Joshua Guedalia , Shaltiel Shmidman , Moshe Koppel , Reut Tsarfaty

For languages with simple morphology, such as English, automatic annotation pipelines such as spaCy or Stanford's CoreNLP successfully serve projects in academia and the industry. For many morphologically-rich languages (MRLs), similar…

Computation and Language · Computer Science 2019-08-16 Reut Tsarfaty , Amit Seker , Shoval Sadde , Stav Klein

Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no parallel text…

Computation and Language · Computer Science 2025-08-05 Motaz Saad , David Langlois , Kamel Smaili

This study presents the development and evaluation of a ByT5-based multilingual translation model tailored for translating the Bible into underrepresented languages. Utilizing the comprehensive Johns Hopkins University Bible Corpus, we…

Computation and Language · Computer Science 2024-06-03 Corinne Aars , Lauren Adams , Xiaokan Tian , Zhaoyu Wang , Colton Wismer , Jason Wu , Pablo Rivas , Korn Sooksatra , Matthew Fendt

In this paper, we present a quantitative evaluation of differences between alternative translations in a large recently released Finnish paraphrase corpus focusing in particular on non-trivial variation in translation. We combine a series…

Computation and Language · Computer Science 2021-05-07 Li-Hsin Chang , Sampo Pyysalo , Jenna Kanerva , Filip Ginter

Current benchmarks for Hebrew Natural Language Processing (NLP) focus mainly on morpho-syntactic tasks, neglecting the semantic dimension of language understanding. To bridge this gap, we set out to deliver a Hebrew Machine Reading…

Computation and Language · Computer Science 2025-08-05 Amir DN Cohen , Hilla Merhav , Yoav Goldberg , Reut Tsarfaty

"Leichte Sprache", the German counterpart to Simple English, is a regulated language aiming to facilitate complex written language that would otherwise stay inaccessible to different groups of people. We present a new sentence-aligned…

Computation and Language · Computer Science 2023-05-29 Vanessa Toborek , Moritz Busch , Malte Boßert , Christian Bauckhage , Pascal Welke

Identifying parallel passages in biblical Hebrew (BH) is central to biblical scholarship for understanding intertextual relationships. Traditional methods rely on manual comparison, a labor-intensive process prone to human error. This study…

Computation and Language · Computer Science 2025-07-02 David M. Smiley

We describe a set of bilingual English--French and English--German parallel corpora in which the direction of translation is accurately and reliably annotated. The corpora are diverse, consisting of parliamentary proceedings, literary…

Computation and Language · Computer Science 2016-03-08 Ella Rabinovich , Shuly Wintner , Ofek Luis Lewinsohn

Machine translation between Arabic and Hebrew has so far been limited by a lack of parallel corpora, despite the political and cultural importance of this language pair. Previous work relied on manually-crafted grammars or pivoting via…

Computation and Language · Computer Science 2016-09-27 Yonatan Belinkov , James Glass

Creating an abridged version of a text involves shortening it while maintaining its linguistic qualities. In this paper, we examine this task from an NLP perspective for the first time. We present a new resource, AbLit, which is derived…

Computation and Language · Computer Science 2023-02-14 Melissa Roemmele , Kyle Shaffer , Katrina Olsen , Yiyi Wang , Steve DeNeefe

Recent works in spoken language translation (SLT) have attempted to build end-to-end speech-to-text translation without using source language transcription during learning or decoding. However, while large quantities of parallel texts (such…

Computation and Language · Computer Science 2018-02-12 Ali Can Kocabiyikoglu , Laurent Besacier , Olivier Kraif

We present the first French partition of the OLDI Seed Corpus, our submission to the WMT 2025 Open Language Data Initiative (OLDI) shared task. We detail its creation process, which involved using multiple machine translation systems and a…

Computation and Language · Computer Science 2025-08-05 Malik Marmonier , Benoît Sagot , Rachel Bawden

In this paper, we present the University of Helsinki submissions to the WMT 2019 shared task on news translation in three language pairs: English-German, English-Finnish and Finnish-English. This year, we focused first on cleaning and…

Computation and Language · Computer Science 2019-06-11 Aarne Talman , Umut Sulubacak , Raúl Vázquez , Yves Scherrer , Sami Virpioja , Alessandro Raganato , Arvi Hurskainen , Jörg Tiedemann

It remains a question that how simultaneous interpretation (SI) data affects simultaneous machine translation (SiMT). Research has been limited due to the lack of a large-scale training corpus. In this work, we aim to fill in the gap by…

Computation and Language · Computer Science 2024-04-02 Jinming Zhao , Yuka Ko , Kosuke Doi , Ryo Fukuda , Katsuhito Sudoh , Satoshi Nakamura
‹ Prev 1 2 3 10 Next ›