中文
相关论文

相关论文: HELFI: a Hebrew-Greek-Finnish Parallel Bible Corpu…

200 篇论文

The Universal Morphology UniMorph project is a collaborative effort to improve how NLP handles complex morphology across the world's languages. The project releases annotated morphological data using a universal tagset, the UniMorph schema.…

For language documentation initiatives, transcription is an expensive resource: one minute of audio is estimated to take one hour and a half on average of a linguist's work (Austin and Sallabank, 2013). Recently, collecting aligned…

计算与语言 · 计算机科学 2019-10-14 Marcely Zanon Boito , Aline Villavicencio , Laurent Besacier

Parallel corpora are a valuable resource for machine translation, but at present their availability and utility is limited by genre- and domain-specificity, licensing restrictions, and the basic difficulty of locating parallel texts in all…

cmp-lg · 计算机科学 2007-05-23 Philip Resnik

Text simplification plays a crucial role in improving the accessibility and comprehensibility of written information for diverse audiences, including language learners and readers with limited literacy. Despite its importance, large-scale,…

计算与语言 · 计算机科学 2026-05-12 Kenji Hilasaca , Nouran Khallaf , Serge Sharoff

While supervised fine-tuning (SFT) has been a straightforward approach for tailoring the output of foundation large language model (LLM) to specific preferences, concerns have been raised about the depth of this alignment, with some…

计算与语言 · 计算机科学 2024-04-26 Runzhe Zhan , Xinyi Yang , Derek F. Wong , Lidia S. Chao , Yue Zhang

Advancements in sign language processing have been hindered by a lack of sufficient data, impeding progress in recognition, translation, and production tasks. The absence of comprehensive sign language datasets across the world's sign…

计算与语言 · 计算机科学 2023-11-20 Shester Gueuwou , Sophie Siake , Colin Leong , Mathias Müller

While large language models (LLMs) excel in various natural language tasks in English, their performance in lower-resourced languages like Hebrew, especially for generative tasks such as abstractive summarization, remains unclear. The high…

计算与语言 · 计算机科学 2025-07-14 Tzuf Paz-Argaman , Itai Mondshine , Asaf Achi Mordechai , Reut Tsarfaty

Text alignment finds application in tasks such as citation recommendation and plagiarism detection. Existing alignment methods operate at a single, predefined level and cannot learn to align texts at, for example, sentence and document…

计算与语言 · 计算机科学 2020-10-06 Xuhui Zhou , Nikolaos Pappas , Noah A. Smith

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Authorship attribution asks whether two pieces of text share a writer, but topical confound makes the task deceptively easy: two authors covering the same topic may look more alike than one author covering two topics. Scholarly prose offers…

数字图书馆 · 计算机科学 2026-05-26 Francis Kulumba , Wissam Antoun , Guillaume Vimont , Laurent Romary , Florian Cafiero

While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers. Yet, such corpora are necessary to fully represent the diversity of…

计算与语言 · 计算机科学 2026-04-28 Arthur Amalvy , Vincent Labatut , Xavier Bost , Hen-Hsen Huang

Many European languages possess rich biblical translation histories, yet existing corpora - in prioritizing linguistic breadth - often fail to capture this depth. To address this gap, we introduce a multilingual corpus of 651 New Testament…

计算与语言 · 计算机科学 2026-05-14 Maciej Rapacz , Aleksander Smywiński-Pohl

The Scielo database is an important source of scientific information in Latin America, containing articles from several research domains. A striking characteristic of Scielo is that many of its full-text contents are presented in more than…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Viviane Pereira Moreira , Karin Becker

The linguistic diversity of India poses significant machine translation challenges, especially for underrepresented tribal languages like Bhili, which lack high-quality linguistic resources. This paper addresses the gap by introducing…

计算与语言 · 计算机科学 2025-11-04 Pooja Singh , Shashwat Bhardwaj , Vaibhav Sharma , Sandeep Kumar

While alignment of texts on the sentential level is often seen as being too coarse, and word alignment as being too fine-grained, bi- or multilingual texts which are aligned on a level in-between are a useful resource for many purposes.…

计算与语言 · 计算机科学 2007-05-23 Lea Cyrus , Hendrik Feddes

We present a free Japanese-French parallel corpus. It includes 15M aligned segments and is obtained by compiling and filtering several existing resources. In this paper, we describe the existing resources, their quantity and quality, the…

计算与语言 · 计算机科学 2022-08-30 Raoul Blin , Fabien Cromières

Morphological tasks use large multi-lingual datasets that organize words into inflection tables, which then serve as training and evaluation data for various tasks. However, a closer inspection of these data reveals profound…

计算与语言 · 计算机科学 2022-10-20 Omer Goldman , Reut Tsarfaty

Despite the strong research interest in document-level Machine Translation (MT), the test sets dedicated to this task are still scarce. The existing test sets mainly cover topics from the general domain and fall short on specialised…

计算与语言 · 计算机科学 2025-02-06 Mariam Nakhlé , Marco Dinarelli , Raheel Qader , Emmanuelle Esperança-Rodier , Hervé Blanchon

Document-level natural language inference (DOCNLI) is a new challenging task in natural language processing, aiming at judging the entailment relationship between a pair of hypothesis and premise documents. Current datasets and baselines…

计算与语言 · 计算机科学 2022-10-25 Hao Wang , Yixin Cao , Yangguang Li , Zhen Huang , Kun Wang , Jing Shao

Parallel corpus is a critical resource in machine learning-based translation. The task of collecting, extracting, and aligning texts in order to build an acceptable corpus for doing the translation is very tedious most especially for…

计算与语言 · 计算机科学 2021-11-01 Kristine Mae Adlaon , Nelson Marcos