English
Related papers

Related papers: German Parliamentary Corpus (GerParCor)

200 papers

While multi-party conversations are often less structured than monologues and documents, they are implicitly organized by semantic level correlations across the interactive turns, and dialogue discourse analysis can be applied to predict…

Computation and Language · Computer Science 2021-10-12 Zhengyuan Liu , Nancy F. Chen

The increasing volume of online discussions requires advanced automatic content moderation to maintain responsible discourse. While hate speech detection on social media is well-studied, research on German-language newspaper forums remains…

Computation and Language · Computer Science 2025-05-28 Felix Krejca , Tobias Kietreiber , Alexander Buchelt , Sebastian Neumaier

Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to transcribe Italian parliamentary speeches have relied on…

Digital Libraries · Computer Science 2026-05-21 Luigi Curini , Alfio Ferrara , Giovanni Pagano , Sergio Picascia

The recent development of large language models (LLMs) has spurred discussions about whether LLM-generated "synthetic samples" could complement or replace traditional surveys, considering their training data potentially reflects attitudes…

Artificial Intelligence · Computer Science 2025-04-29 Leah von der Heyde , Anna-Carolina Haensch , Alexander Wenz

Axial coding is a commonly used qualitative analysis method that enhances document understanding by organizing sentence-level open codes into broader categories. In this paper, we operationalize axial coding with large language models…

Computation and Language · Computer Science 2026-01-23 Angelina Parfenova , David Graus , Juergen Pfeffer

Coalition negotiations are a cornerstone of parliamentary democracies, characterised by complex interactions and strategic communications among political parties. Despite its significance, the modelling of these negotiations has remained…

Computation and Language · Computer Science 2024-02-20 Farhad Moghimifar , Yuan-Fang Li , Robert Thomson , Gholamreza Haffari

This paper briefly reports our ongoing attempt at the development of a multi-platform browser-based speech recording system. We designed the system toward a service of providing open service of building large-scale speech corpora at a…

Human-Computer Interaction · Computer Science 2019-12-20 Keita Ishizuka , Takashi Nose

In this paper, we apply transformer-based Natural Language Generation (NLG) techniques to the problem of text simplification. Currently, there are only a few German datasets available for text simplification, even fewer with larger and…

Computation and Language · Computer Science 2023-12-18 Thorben Schomacker , Tillmann Dönicke , Marina Tropmann-Frick

CERN, the European Organization for Nuclear Research, has operated an institutional preprint repository for more than 10 years. The repository contains over 850,000 records of which more than 450,000 are full-text OA preprints, mostly in…

Digital Libraries · Computer Science 2007-05-23 Alberto Pepe , Joanne Yeomans

Speech processing and translation technology have the potential to facilitate meetings of individuals who do not share any common language. To evaluate automatic systems for such a task, a versatile and realistic evaluation corpus is…

Computation and Language · Computer Science 2025-12-24 Marko Čechovič , Natália Komorníková , Dominik Macháček , Ondřej Bojar

Large datasets are essential for neural modeling of many NLP tasks. Current publicly available open-domain dialogue datasets offer a trade-off between quality (e.g., DailyDialog) and size (e.g., Opensubtitles). We narrow this gap by…

Computation and Language · Computer Science 2021-01-25 Richard Csaky , Gabor Recski

Organisations disclose their privacy practices by posting privacy policies on their website. Even though users often care about their digital privacy, they often don't read privacy policies since they require a significant investment in…

Information Retrieval · Computer Science 2024-04-02 Mukund Srinath , Shomir Wilson , C. Lee Giles

We present Knesset-DictaBERT, a large Hebrew language model fine-tuned on the Knesset Corpus, which comprises Israeli parliamentary proceedings. The model is based on the DictaBERT architecture and demonstrates significant improvements in…

Computation and Language · Computer Science 2024-07-31 Gili Goldin , Shuly Wintner

In this paper, we present CORE-GPT, a novel question-answering platform that combines GPT-based language models and more than 32 million full-text open access scientific articles from CORE. We first demonstrate that GPT3.5 and GPT4 cannot…

Computation and Language · Computer Science 2023-07-11 David Pride , Matteo Cancellieri , Petr Knoth

While the progress of machine translation of written text has come far in the past several years thanks to the increasing availability of parallel corpora and corpora-based training technologies, automatic translation of spoken text and…

Computation and Language · Computer Science 2020-08-06 Matīss Rikters , Ryokan Ri , Tong Li , Toshiaki Nakazawa

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

It is now a common practice to compare models of human language processing by predicting participant reactions (such as reading times) to corpora consisting of rich naturalistic linguistic materials. However, many of the corpora used in…

Computation and Language · Computer Science 2017-08-22 Richard Futrell , Edward Gibson , Hal Tily , Idan Blank , Anastasia Vishnevetsky , Steven T. Piantadosi , Evelina Fedorenko

Moralizations - arguments that invoke moral values to justify demands or positions - are a yet underexplored form of persuasive communication. We present the Moralization Corpus, a novel multi-genre dataset designed to analyze how moral…

Computation and Language · Computer Science 2026-03-19 Maria Becker , Mirko Sommer , Lars Tapken , Yi Wan Teh , Bruno Brocai

Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obtaining such data from…

Computation and Language · Computer Science 2016-03-23 Krzysztof Wołk , Emilia Rejmund , Krzysztof Marasek

Recent approaches to text analysis from social media and other corpora rely on word lists to detect topics, measure meaning, or to select relevant documents. These lists are often generated by applying computational lexicon expansion…

Computation and Language · Computer Science 2023-03-14 Anna Di Natale , David Garcia