English
Related papers

Related papers: Compiling and Processing Historical and Contempora…

200 papers

Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded…

The first step in discourse analysis involves dividing a text into segments. We annotate the first high-quality small-scale medical corpus in English with discourse segments and analyze how well news-trained segmenters perform on this…

Computation and Language · Computer Science 2019-04-16 Elisa Ferracane , Titan Page , Junyi Jessy Li , Katrin Erk

Pre-training large-scale language models (LMs) requires huge amounts of text corpora. LMs for English enjoy ever growing corpora of diverse language resources. However, less resourced languages and their mono- and multilingual LMs often…

Computation and Language · Computer Science 2020-07-07 Maria Khvalchik , Mikhail Galkin

Large-scale pretrained language models have become ubiquitous in Natural Language Processing. However, most of these models are available either in high-resource languages, in particular English, or as multilingual models that compromise…

Computation and Language · Computer Science 2020-09-21 Stefan Daniel Dumitrescu , Andrei-Marius Avram , Sampo Pyysalo

Old French is a typical example of an under-resourced historic languages, that furtherly displays animportant amount of linguistic variation. In this paper, we present the current results of a long going project (2015-...) and describe how…

Computation and Language · Computer Science 2021-09-24 Jean-Baptiste Camps , Thibault Clérice , Frédéric Duval , Lucence Ing , Naomi Kanaoka , Ariane Pinche

This article presents two corpora of English and Czech texts generated with large language models (LLMs). The motivation is to create a resource for comparing human-written texts with LLM-generated text linguistically. Emphasis was placed…

Computation and Language · Computer Science 2025-11-11 Jiří Milička , Anna Marklová , Václav Cvrček

Lecture transcript translation helps learners understand online courses, however, building a high-quality lecture machine translation system lacks publicly available parallel corpora. To address this, we examine a framework for parallel…

Computation and Language · Computer Science 2023-11-08 Haiyue Song , Raj Dabre , Chenhui Chu , Atsushi Fujita , Sadao Kurohashi

Most well-established data collection methods currently adopted in NLP depend on the assumption of speaker literacy. Consequently, the collected corpora largely fail to represent swathes of the global population, which tend to be some of…

Computation and Language · Computer Science 2021-02-08 Stephanie Hirmer , Alycia Leonard , Josephine Tumwesige , Costanza Conforti

A child's spoken ability continues to change until their adult age. Until 7-8yrs, their speech sound development and language structure evolve rapidly. This dynamic shift in their spoken communication skills and data privacy make it…

Sound · Computer Science 2025-07-18 John Hansen , Satwik Dutta , Ellen Grand

With the increasing popularity of mobile devices and the wide adoption of mobile Apps, an increasing concern of privacy issues is raised. Privacy policy is identified as a proper medium to indicate the legal terms, such as GDPR, and to bind…

Computers and Society · Computer Science 2020-05-15 Shuang Liu , Renjie Guo , Baiyang Zhao , Tao Chen , Meishan Zhang

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian Portuguese…

Information Retrieval · Computer Science 2024-04-11 Mirelle Bueno , Eduardo Seiti de Oliveira , Rodrigo Nogueira , Roberto A. Lotufo , Jayr Alencar Pereira

We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-20th centuries. Finding aids are handwritten documents that…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Solène Tarride , Mélodie Boillet , Jean-François Moufflet , Christopher Kermorvant

In this paper, we aim at the application of Natural Language Processing (NLP) techniques to historical research endeavors, particularly addressing the study of religious invectives in the context of the Protestant Reformation in Tudor…

Computation and Language · Computer Science 2025-09-29 Sophie Spliethoff , Sanne Hoeken , Silke Schwandt , Sina Zarrieß , Özge Alaçam

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particular non-standardized…

Computation and Language · Computer Science 2023-04-20 Verena Blaschke , Hinrich Schütze , Barbara Plank

This short paper examines diagrams describing neural network systems in academic conference proceedings. Many aspects of scholarly communication are controlled, particularly with relation to text and formatting, but often diagrams are not…

Human-Computer Interaction · Computer Science 2021-05-03 Guy Clarke Marshall , Caroline Jay , Andre Freitas

This paper presents a practical pipeline for turning text corpora into quantitative semantic signals. Each news item is represented as a full-document embedding, scored through logprob-based evaluation over a configurable positional…

Computation and Language · Computer Science 2026-04-16 Hugo Moreira

Previous work in the social sciences, psychology and linguistics has show that liars have some control over the content of their stories, however their underlying state of mind may "leak out" through the way that they tell them. To the best…

Computation and Language · Computer Science 2021-04-05 Francielle Alves Vargas , Thiago Alexandre Salgueiro Pardo

Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal…

Computation and Language · Computer Science 2025-06-09 Stefanie Urchs , Veronika Thurner , Matthias Aßenmacher , Christian Heumann , Stephanie Thiemichen

Organisations disclose their privacy practices by posting privacy policies on their website. Even though users often care about their digital privacy, they often don't read privacy policies since they require a significant investment in…

Information Retrieval · Computer Science 2024-04-02 Mukund Srinath , Shomir Wilson , C. Lee Giles

Being able to understand information is a key factor for a self-determined life and society. It is also very important for participating in democratic processes. The study of automatic text simplification is often limited by the…

Computation and Language · Computer Science 2026-03-17 Stefan Bott , Verena Riegler , Horacio Saggion , Almudena Rascón Alcaina , Nouran Khallaf
‹ Prev 1 8 9 10 Next ›