中文
相关论文

相关论文: Compiling and Processing Historical and Contempora…

200 篇论文

Although there are increasing and significant ties between China and Portuguese-speaking countries, there is not much parallel corpora in the Chinese-Portuguese language pair. Both languages are very populous, with 1.2 billion native…

计算与语言 · 计算机科学 2018-04-06 Siyou Liu , Longyue Wang , Chao-Hong Liu

In natural language processing (NLP), there is a need for more resources in Portuguese, since much of the data used in the state-of-the-art research is in other languages. In this paper, we pretrain a T5 model on the BrWac corpus, an…

计算与语言 · 计算机科学 2020-10-12 Diedre Carmo , Marcos Piau , Israel Campiotti , Rodrigo Nogueira , Roberto Lotufo

High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder based on the…

Historical Document Processing is the process of digitizing written material from the past for future use by historians and other scholars. It incorporates algorithms and software tools from various subfields of computer science, including…

计算机视觉与模式识别 · 计算机科学 2020-09-14 James P. Philips , Nasseh Tabrizi

The objective of this paper is to develop predictive models to classify Brazilian legal proceedings in three possible classes of status: (i) archived proceedings, (ii) active proceedings, and (iii) suspended proceedings. This problem's…

计算与语言 · 计算机科学 2021-06-24 Felipe Maia Polo , Itamar Ciochetti , Emerson Bertolo

In this paper we present PeLLE, a family of large language models based on the RoBERTa architecture, for Brazilian Portuguese, trained on curated, open data from the Carolina corpus. Aiming at reproducible results, we describe details of…

This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata…

计算与语言 · 计算机科学 2025-04-18 Sergio Torres Aguilar

We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions…

计算与语言 · 计算机科学 2018-06-13 Benjamin Nye , Junyi Jessy Li , Roma Patel , Yinfei Yang , Iain J. Marshall , Ani Nenkova , Byron C. Wallace

Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of…

计算与语言 · 计算机科学 2025-07-25 Nicholas Kluge Corrêa , Aniket Sen , Sophia Falk , Shiza Fatimah

This paper reports on the preliminary phase of our ongoing research towards developing an intelligent tutoring environment for Turkish grammar. One of the components of this environment is a corpus search tool which, among other aspects of…

cmp-lg · 计算机科学 2016-08-31 H. Altay Guvenir , Kemal Oflazer

Historical newspapers are a source of research for the human and social sciences. However, these image collections are difficult to read by machine due to the low quality of the print, the lack of standardization of the pages in addition to…

信息检索 · 计算机科学 2020-02-21 José E. B. Maia , Gildácio J. de A. Sá

While the progress of machine translation of written text has come far in the past several years thanks to the increasing availability of parallel corpora and corpora-based training technologies, automatic translation of spoken text and…

计算与语言 · 计算机科学 2020-08-06 Matīss Rikters , Ryokan Ri , Tong Li , Toshiaki Nakazawa

Although Natural Language Processing (NLP) research on argument mining has advanced considerably in recent years, most studies draw on corpora of asynchronous and written texts, often produced by individuals. Few published corpora of…

计算与语言 · 计算机科学 2020-05-26 Christopher Olshefski , Luca Lugini , Ravneet Singh , Diane Litman , Amanda Godley

Insightful findings in political science often require researchers to analyze documents of a certain subject or type, yet these documents are usually contained in large corpora that do not distinguish between pertinent and non-pertinent…

计算与语言 · 计算机科学 2019-10-29 Shrey Desai , Barea Sinno , Alex Rosenfeld , Junyi Jessy Li

In this work, we develop a pipeline for historical-psychological text analysis in classical Chinese. Humans have produced texts in various languages for thousands of years; however, most of the computational literature is focused on…

计算与语言 · 计算机科学 2025-04-17 Yuqi Chen , Sixuan Li , Ying Li , Mohammad Atari

The need for raw large raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to…

计算与语言 · 计算机科学 2022-01-19 Julien Abadji , Pedro Ortiz Suarez , Laurent Romary , Benoît Sagot

Research in Computational Linguistics is dependent on text corpora for training and testing new tools and methodologies. While there exists a plethora of annotated linguistic information, these corpora are often not interoperable without…

计算与语言 · 计算机科学 2020-11-03 Timo Lek , Anna de Groot , Tobias Kuhn , Roser Morante

In this paper we describe an architecture and functionality of main components of a workbench for an acquisition of domain knowledge from large text corpora. The workbench supports an incremental process of corpus analysis starting from a…

cmp-lg · 计算机科学 2008-02-03 Andrei Mikheev , Steven Finch

This paper investigates the impact of corpus creation decisions on large multi-lingual geographic web corpora. Beginning with a 427 billion word corpus derived from the Common Crawl, three methods are used to improve the quality of…

计算与语言 · 计算机科学 2024-03-14 Jonathan Dunn

The accelerated dissemination of disinformation often outpaces the capacity for manual fact-checking, highlighting the urgent need for Semi-Automated Fact-Checking (SAFC) systems. Within the Portuguese language context, there is a noted…

计算与语言 · 计算机科学 2025-08-12 Juliana Resplande Sant'anna Gomes , Arlindo Rodrigues Galvão Filho