中文
相关论文

相关论文: HELFI: a Hebrew-Greek-Finnish Parallel Bible Corpu…

200 篇论文

Technology has increasingly become an integral part of the Bible translation process. Over time, both the translation process and relevant technology have evolved greatly. More recently, the field of Natural Language Processing (NLP) has…

计算与语言 · 计算机科学 2023-02-03 Joel Mathew , Ulf Hermjakob

Measuring the semantic similarity of different texts has many important applications in Digital Humanities research such as information retrieval, document clustering and text summarization. The performance of different methods depends on…

计算与语言 · 计算机科学 2021-12-24 Quan Duong , Mika Hämäläinen , Khalid Alnajjar

DEMorphy is a morphological analyzer for German. It is built onto large, compactified lexicons from German Morphological Dictionary. A guesser based on German declension suffixed is also provided. For German, we provided a state-of-art…

计算与语言 · 计算机科学 2018-03-05 Duygu Altinok

The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME (Biblioteca Regional de Medicina) in agreement with the Pan American Health…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Martin Krallinger

Quality estimation (QE) plays a crucial role in machine translation (MT) workflows, as it serves to evaluate generated outputs that have no reference translations and to determine whether human post-editing or full retranslation is…

计算与语言 · 计算机科学 2026-03-13 Assaf Siani , Anna Kernerman , Ilan Kernerman

In this paper we present Morphy, an integrated tool for German morphology, part-of-speech tagging and context-sensitive lemmatization. Its large lexicon of more than 320,000 word forms plus its ability to process German compound nouns…

计算与语言 · 计算机科学 2007-05-23 Wolfgang Lezius , Reinhard Rapp , Manfred Wettler

We present a semantic role labeling resource for Hebrew built semi-automatically through annotation projection from English. This corpus is derived from the multilingual OpenSubtitles dataset and includes short informal sentences, for which…

计算与语言 · 计算机科学 2020-05-19 Ben Eyal , Michael Elhadad

Recently unsupervised Bilingual Lexicon Induction (BLI) without any parallel corpus has attracted much research interest. One of the crucial parts in methods for the BLI task is the matching procedure. Previous works impose a too strong…

计算与语言 · 计算机科学 2020-10-15 Xu Zhao , Zihao Wang , Hao Wu , Yong Zhang

Prior studies in multilingual language modeling (e.g., Cotterell et al., 2018; Mielke et al., 2019) disagree on whether or not inflectional morphology makes languages harder to model. We attempt to resolve the disagreement and extend those…

计算与语言 · 计算机科学 2021-03-29 Hyunji Hayley Park , Katherine J. Zhang , Coleman Haley , Kenneth Steimel , Han Liu , Lane Schwartz

We describe a simple procedure for the automatic creation of word-level alignments between printed documents and their respective full-text versions. The procedure is unsupervised, uses standard, off-the-shelf components only, and reaches…

计算与语言 · 计算机科学 2021-05-03 Mark-Christoph Müller , Sucheta Ghosh , Ulrike Wittig , Maja Rey

We translate a closed text that is known in advance and available in many languages into a new and severely low resource language. Most human translation efforts adopt a portion-based approach to translate consecutive pages/chapters in…

计算与语言 · 计算机科学 2021-10-28 Zhong Zhou , Alex Waibel

Machine Translation (MT) for Ancient Greek (AG) to Modern Greek (MG) is a low-resource task, constrained by the lack of large-scale, high-quality parallel data. We address this gap by introducing the AG-MG Parallel Corpus, a new resource…

计算与语言 · 计算机科学 2026-05-19 Spyridon Mavromatis , Sokratis Sofianopoulos , Prokopis Prokopidis , Maria Giagkou

Multilingual data from the web is essential for LLM pretraining. Yet, scraping it is expensive, and research groups repeatedly crawl the same content. For example, we found that over 40\% of tokens across major Arabic web corpora are…

计算与语言 · 计算机科学 2026-01-30 Sultan Alrashed , Francesco Orabona

We introduce MorphNLI, a modular step-by-step approach to natural language inference (NLI). When classifying the premise-hypothesis pairs into {entailment, contradiction, neutral}, we use a language model to generate the necessary edits to…

计算与语言 · 计算机科学 2026-02-16 Vlad Andrei Negru , Robert Vacareanu , Camelia Lemnaru , Mihai Surdeanu , Rodica Potolea

This article presents two corpora of English and Czech texts generated with large language models (LLMs). The motivation is to create a resource for comparing human-written texts with LLM-generated text linguistically. Emphasis was placed…

计算与语言 · 计算机科学 2025-11-11 Jiří Milička , Anna Marklová , Václav Cvrček

We train a bilingual Arabic-Hebrew language model using a transliterated version of Arabic texts in Hebrew, to ensure both languages are represented in the same script. Given the morphological, structural similarities, and the extensive…

计算与语言 · 计算机科学 2024-02-27 Aviad Rom , Kfir Bar

There are a lot of tools and resources available for processing Finnish. In this paper, we survey recent papers focusing on Finnish NLP related to many different subcategories of NLP such as parsing, generation, semantics and speech. NLP…

计算与语言 · 计算机科学 2021-09-24 Mika Hämäläinen , Khalid Alnajjar

Text simplification is an intralingual translation task in which documents, or sentences of a complex source text are simplified for a target audience. The success of automatic text simplification systems is highly dependent on the quality…

计算与语言 · 计算机科学 2024-07-08 Regina Stodden , Omar Momen , Laura Kallmeyer

We propose a language-independent approach for improving statistical machine translation for morphologically rich languages using a hybrid morpheme-word representation where the basic unit of translation is the morpheme, but word boundaries…

计算与语言 · 计算机科学 2019-11-20 Minh-Thang Luong , Preslav Nakov , Min-Yen Kan

Bilingual annotators were paid to link roughly sixteen thousand corresponding words between on-line versions of the Bible in modern French and modern English. These annotations are freely available to the research community from…

cmp-lg · 计算机科学 2007-05-23 I. Dan Melamed