English
Related papers

Related papers: Compiling and Processing Historical and Contempora…

200 papers

Large parallel corpora that are automatically obtained from the web, documents or elsewhere often exhibit many corrupted parts that are bound to negatively affect the quality of the systems and models that learn from these corpora. This…

Computation and Language · Computer Science 2018-10-22 Matīss Rikters

Mathematics is a highly specialized domain with its own unique set of challenges. Despite this, there has been relatively little research on natural language processing for mathematical texts, and there are few mathematical language…

Computation and Language · Computer Science 2024-06-18 Jacob Collard , Valeria de Paiva , Eswaran Subrahmanian

This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Brazilian Chamber of Deputies, paired with news articles and…

Computation and Language · Computer Science 2025-08-25 Leandro Carísio Fernandes , Guilherme Zeferino Rodrigues Dobins , Roberto Lotufo , Jayr Alencar Pereira

At a time when the quantity of - more or less freely - available data is increasing significantly, thanks to digital corpora, editions or libraries, the development of data mining tools or deep learning methods allows researchers to build a…

Computer Vision and Pattern Recognition · Computer Science 2019-04-29 Jean-Baptiste Camps , Gilles Guilhem Couffignal

This paper presents a new annotated corpus of 513 anonymized radiology reports written in Spanish. Reports were manually annotated with entities, negation and uncertainty terms and relations. The corpus was conceived as an evaluation…

Computation and Language · Computer Science 2017-11-01 Viviana Cotik , Darío Filippo , Roland Roller , Hans Uszkoreit , Feiyu Xu

The goal of this project is to (i) accumulate annotated informal/formal mathematical corpora suitable for training semi-automated translation between informal and formal mathematics by statistical machine-translation methods, (ii) to…

Artificial Intelligence · Computer Science 2014-05-15 Cezary Kaliszyk , Josef Urban , Jiri Vyskocil , Herman Geuvers

In a representative democracy, some decide in the name of the rest, and these elected officials are commonly gathered in public assemblies, such as parliaments, where they discuss policies, legislate, and vote on fundamental initiatives. A…

Computation and Language · Computer Science 2022-07-01 Paulo Almeida , Manuel Marques-Pita , Joana Gonçalves-Sá

To advance the neural encoding of Portuguese (PT), and a fortiori the technological preparation of this language for the digital age, we developed a Transformer-based foundation model that sets a new state of the art in this respect for two…

Computation and Language · Computer Science 2024-04-03 João Rodrigues , Luís Gomes , João Silva , António Branco , Rodrigo Santos , Henrique Lopes Cardoso , Tomás Osório

Leveraging research on the neural modelling of Portuguese, we contribute a collection of datasets for an array of language processing tasks and a corresponding collection of fine-tuned neural language models on these downstream tasks. To…

Computation and Language · Computer Science 2024-05-10 Tomás Osório , Bernardo Leite , Henrique Lopes Cardoso , Luís Gomes , João Rodrigues , Rodrigo Santos , António Branco

We survey clinical document corpora, with focus on German textual data. Due to rigid data privacy legislation in Germany these resources, with only few exceptions, are stored in safe clinical data spaces and locked against clinic-external…

Computation and Language · Computer Science 2025-02-20 Udo Hahn

The article examines the theoretical, methodological, and technical foundations of research on audiovisual corpora within the field of digital humanities. It outlines the main transversal issues underlying the processes of constructing,…

Digital Libraries · Computer Science 2025-11-07 Peter Stockinger

Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of…

Computation and Language · Computer Science 2025-01-22 Marlo Souza , Bruno Cabral , Daniela Claro , Lais Salvador

In the prose style transfer task a system, provided with text input and a target prose style, produces output which preserves the meaning of the input text but alters the style. These systems require parallel data for evaluation of results…

Computation and Language · Computer Science 2021-09-01 Keith Carlson , Allen Riddell , Daniel Rockmore

Researchers working in areas such as lexicography, translation studies, and computational linguistics, use a combination of automated and semi-automated tools to analyze the content of text corpora. Keywords, named entities, and events are…

Human-Computer Interaction · Computer Science 2022-03-24 Shane Sheehan , Saturnino Luz , Masood Masoodian

Most previous work on the recently developed language-modeling approach to information retrieval focuses on document-specific characteristics, and therefore does not take into account the structure of the surrounding corpus. We propose a…

Information Retrieval · Computer Science 2007-05-23 Oren Kurland , Lillian Lee

The majority of NLG systems have been designed following either a template-based or a pipeline-based architecture. Recent neural models for data-to-text generation have been proposed with an end-to-end deep learning flavor, which handles…

Computation and Language · Computer Science 2022-10-11 Yan V. Sym , João Gabriel M. Campos , Marcos M. José , Fabio G. Cozman

Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obtaining such data from…

Computation and Language · Computer Science 2016-03-23 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties…

Computation and Language · Computer Science 2022-05-11 Andreas Liesenfeld , Mark Dingemanse

In populous countries, pending legal cases have been growing exponentially. There is a need for developing techniques for processing and organizing legal documents. In this paper, we introduce a new corpus for structuring legal documents.…

Computation and Language · Computer Science 2022-09-20 Prathamesh Kalamkar , Aman Tiwari , Astha Agarwal , Saurabh Karn , Smita Gupta , Vivek Raghavan , Ashutosh Modi

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in…

Computation and Language · Computer Science 2022-07-04 Asier Gutiérrez-Fandiño , David Pérez-Fernández , Jordi Armengol-Estapé , David Griol , Zoraida Callejas