中文
相关论文

相关论文: Compiling and Processing Historical and Contempora…

200 篇论文

This paper presents the first publicly available version of the Carolina Corpus and discusses its future directions. Carolina is a large open corpus of Brazilian Portuguese texts under construction using web-as-corpus methodology enhanced…

This paper presents a number of experiments to model changes in a historical Portuguese corpus composed of literary texts for the purpose of temporal text classification. Algorithms were trained to classify texts with respect to their…

计算与语言 · 计算机科学 2016-10-04 Marcos Zampieri , Shervin Malmasi , Mark Dras

In Brazil, the governmental body responsible for overseeing and coordinating post-graduate programs, CAPES, keeps records of all theses and dissertations presented in the country. Information regarding such documents can be accessed online…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Gabrielli Harumi Yamashita , Michel Jose Anzanello

We describe a set of bilingual English--French and English--German parallel corpora in which the direction of translation is accurately and reliably annotated. The corpora are diverse, consisting of parliamentary proceedings, literary…

计算与语言 · 计算机科学 2016-03-08 Ella Rabinovich , Shuly Wintner , Ofek Luis Lewinsohn

The performance of large language models (LLMs) is deeply influenced by the quality and composition of their training data. While much of the existing work has centered on English, there remains a gap in understanding how to construct…

计算与语言 · 计算机科学 2025-09-11 Thales Sales Almeida , Rodrigo Nogueira , Helio Pedrini

This paper will present textual corpora for Serbian (and Serbo-Croatian), usable for the training of large language models and publicly available at one of the several notable online repositories. Each corpus will be classified using…

计算与语言 · 计算机科学 2024-05-16 Mihailo Škorić , Nikola Janković

This paper presents two significant contributions: First, it introduces a novel dataset of 19th-century Latin American newspaper texts, addressing a critical gap in specialized corpora for historical and linguistic analysis in this region.…

计算与语言 · 计算机科学 2025-03-31 Laura Manrique-Gómez , Tony Montes , Arturo Rodríguez-Herrera , Rubén Manrique

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in…

Speech provides a natural way for human-computer interaction. In particular, speech synthesis systems are popular in different applications, such as personal assistants, GPS applications, screen readers and accessibility tools. However, not…

We present a state-of-the-art report on visualization corpora in automated chart analysis research. We survey 56 papers that created or used a visualization corpus as the input of their research techniques or systems. Based on a multi-level…

人机交互 · 计算机科学 2023-08-10 Chen Chen , Zhicheng Liu

This paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained on large corpora…

计算与语言 · 计算机科学 2020-03-16 Serge Sharoff

The Scielo database is an important source of scientific information in Latin America, containing articles from several research domains. A striking characteristic of Scielo is that many of its full-text contents are presented in more than…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Viviane Pereira Moreira , Karin Becker

Sentiment Analysis is one of the most classical and primarily studied natural language processing tasks. This problem had a notable advance with the proposition of more complex and scalable machine learning models. Despite this progress,…

计算与语言 · 计算机科学 2021-12-13 Frederico Souza , João Filho

This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological…

计算与语言 · 计算机科学 2025-12-09 Revekka Kyriakoglou , Anna Pappa

The increasing volume of scientific research necessitates effective communication across language barriers. Machine translation (MT) offers a promising solution for accessing international publications. However, the scientific domain…

计算与语言 · 计算机科学 2026-05-21 Dimitris Roussis , Sokratis Sofianopoulos , Stelios Piperidis

Information about pretraining corpora used to train the current best-performing language models is seldom discussed: commercial models rarely detail their data, and even open models are often released without accompanying training data or…

Textual health records of cancer patients are usually protracted and highly unstructured, making it very time-consuming for health professionals to get a complete overview of the patient's therapeutic course. As such limitations can lead to…

计算与语言 · 计算机科学 2023-08-08 Hugo Sousa , Arian Pasquali , Alípio Jorge , Catarina Sousa Santos , Mário Amorim Lopes

In this paper we describe the Portuguese-language podcast dataset we have released for academic research purposes. We give an overview of how the data was sampled, descriptive statistics over the collection, as well as information about the…

In this paper, we propose a study of progressive development of the structure of corpus collected starting from the discussion forums. These are the first results from an ethnographic description of constitution and analysis practices of a…

数字图书馆 · 计算机科学 2018-08-10 Goritsa Ninova , Hassan Atifi

We report on two corpora to be used in the evaluation of component systems for the tasks of (1) linear segmentation of text and (2) summary-directed sentence extraction. We present characteristics of the corpora, methods used in the…

计算与语言 · 计算机科学 2007-05-23 Judith L. Klavans , Kathleen R. McKeown , Min-Yen Kan , Susan Lee
‹ 上一页 1 2 3 10 下一页 ›