中文
相关论文

相关论文: Compiling and Processing Historical and Contempora…

200 篇论文

Automatic terminology processing appeared 10 years ago when electronic corpora became widely available. Such processing may be statistically or linguistically based and produces terminology resources that can be used in a number of…

计算机与社会 · 计算机科学 2014-12-16 C. Enguehard , B. Daille , E. Morin

This article explores the requirements for corpus compilation within the GiesKaNe project (University of Giessen and Kassel, Syntactic Basic Structures of New High German). The project is defined by three central characteristics: it is a…

计算与语言 · 计算机科学 2025-02-10 Volker Emmrich

Discourse parsing is an integral part of understanding information flow and argumentative structure in documents. Most previous research has focused on inducing and evaluating models from the English RST Discourse Treebank. However,…

计算与语言 · 计算机科学 2017-01-12 Chloé Braud , Maximin Coavoux , Anders Søgaard

Parallel corpora are a valuable resource for machine translation, but at present their availability and utility is limited by genre- and domain-specificity, licensing restrictions, and the basic difficulty of locating parallel texts in all…

cmp-lg · 计算机科学 2007-05-23 Philip Resnik

Text readability assessment has gained significant attention from researchers in various domains. However, the lack of exploration into corpus compatibility poses a challenge as different research groups utilize different corpora. In this…

计算与语言 · 计算机科学 2023-09-14 Zhenzhen Li , Han Ding , Shaohong Zhang

Illustrations are an essential transmission instrument. For an historian, the first step in studying their evolution in a corpus of similar manuscripts is to identify which ones correspond to each other. This image collation task is…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Ryad Kaoua , Xi Shen , Alexandra Durr , Stavros Lazaris , David Picard , Mathieu Aubry

This paper gives comprehensive analyses of corpora based on Wikipedia for several tasks in question answering. Four recent corpora are collected,WikiQA, SelQA, SQuAD, and InfoQA, and first analyzed intrinsically by contextual similarities,…

计算与语言 · 计算机科学 2018-02-06 Tomasz Jurczyk , Amit Deshmane , Jinho D. Choi

This paper presents a grammar formalism designed for use in data-oriented approaches to language processing. The formalism is best described as a right-linear indexed grammar extended in linguistically interesting ways. The paper goes on to…

cmp-lg · 计算机科学 2016-08-31 David Tugwell

This paper presents a semi-automatic approach to create a diachronic corpus of voices balanced for speaker's age, gender, and recording period, according to 32 categories (2 genders, 4 age ranges and 4 recording periods). Corpora were…

音频与语音处理 · 电气工程与系统科学 2024-04-29 Rémi Uro , David Doukhan , Albert Rilliard , Laëtitia Larcher , Anissa-Claire Adgharouamane , Marie Tahon , Antoine Laurent

This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and is freely available…

计算与语言 · 计算机科学 2018-02-01 Pavel Král , Ladislav Lenc

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

计算与语言 · 计算机科学 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

This report provides an overview of the CorCenCC project and the online corpus resource that was developed as a result of work on the project. The report lays out the theoretical underpinnings of the research, demonstrating how the project…

计算与语言 · 计算机科学 2020-10-13 Dawn Knight , Steve Morris , Tess Fitzpatrick , Paul Rayson , Irena Spasić , Enlli Môn Thomas

The objective of the PANACEA ICT-2007.2.2 EU project is to build a platform that automates the stages involved in the acquisition, production, updating and maintenance of the large language resources required by, among others, MT systems.…

计算与语言 · 计算机科学 2013-03-11 Núria Bel , Vassilis Papavasiliou , Prokopis Prokopidis , Antonio Toral , Victoria Arranz

We consider three major text sources about the Tang Dynasty of China in our experiments that aim to segment text written in classical Chinese. These corpora include a collection of Tang Tomb Biographies, the New Tang Book, and the Old Tang…

计算与语言 · 计算机科学 2020-07-23 Chao-Lin Liu , Chang-Ting Chu , Wei-Ting Chang , Ti-Yong Zheng

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

计算与语言 · 计算机科学 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

Applying methods in natural language processing on electronic health records (EHR) data is a growing field. Existing corpus and annotation focus on modeling textual features and relation prediction. However, there is a paucity of annotated…

计算与语言 · 计算机科学 2022-04-08 Yanjun Gao , Dmitriy Dligach , Timothy Miller , Samuel Tesch , Ryan Laffin , Matthew M. Churpek , Majid Afshar

Parliamentary debates represent a large and partly unexploited treasure trove of publicly accessible texts. In the German-speaking area, there is a certain deficit of uniformly accessible and annotated corpora covering all German-speaking…

计算与语言 · 计算机科学 2022-04-25 Giuseppe Abrami , Mevlüt Bagci , Leon Hammerla , Alexander Mehler

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for mining such data from…

计算与语言 · 计算机科学 2015-11-20 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

The impression section of a radiology report summarizes important radiology findings and plays a critical role in communicating these findings to physicians. However, the preparation of these summaries is time-consuming and error-prone for…

Whereas much of the success of the current generation of neural language models has been driven by increasingly large training corpora, relatively little research has been dedicated to analyzing these massive sources of textual data. In…

计算与语言 · 计算机科学 2021-06-02 Alexandra Sasha Luccioni , Joseph D. Viviano