English
Related papers

Related papers: German Parliamentary Corpus (GerParCor)

200 papers

In this work, we present PoliCorp (https://demo-pollux.gesis.org/), a web portal designed to facilitate the search and analysis of political text corpora. PoliCorp provides researchers with access to rich textual data, enabling in-depth…

Digital Libraries · Computer Science 2025-09-23 Nina Smirnova , Muhammad Ahsan Shahid , Philipp Mayr

The lack of publicly accessible text corpora is a major obstacle for progress in natural language processing. For medical applications, unfortunately, all language communities other than English are low-resourced. In this work, we present…

Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This problem is exacerbated for non-English languages, where openly…

Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal…

Computation and Language · Computer Science 2025-06-09 Stefanie Urchs , Veronika Thurner , Matthias Aßenmacher , Christian Heumann , Stephanie Thiemichen

The application of natural language processing on political texts as well as speeches has become increasingly relevant in political sciences due to the ability to analyze large text corpora which cannot be read by a single person. But such…

Computation and Language · Computer Science 2024-10-24 Kai-Robin Lange , Carsten Jentsch

This paper presents a new long-form release of the Swiss Parliaments Corpus, converting entire multi-hour Swiss German debate sessions (each aligned with the official session protocols) into high-quality speech-text pairs. Our pipeline…

Computation and Language · Computer Science 2026-03-13 Vincenzo Timmel , Manfred Vogel , Daniel Perruchoud , Reza Kakooee

Social media serves as a critical medium in modern politics because it both reflects politicians' ideologies and facilitates communication with younger generations. We present MultiParTweet, a multilingual tweet corpus from X that connects…

Computation and Language · Computer Science 2025-12-15 Mevlüt Bagci , Ali Abusaleh , Daniel Baumartz , Giueseppe Abrami , Maxim Konca , Alexander Mehler

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particular non-standardized…

Computation and Language · Computer Science 2023-04-20 Verena Blaschke , Hinrich Schütze , Barbara Plank

We present the Swiss Parliaments Corpus (SPC), an automatically aligned Swiss German speech to Standard German text corpus. This first version of the corpus is based on publicly available data of the Bernese cantonal parliament and consists…

Computation and Language · Computer Science 2021-06-10 Michel Plüss , Lukas Neukom , Christian Scheller , Manfred Vogel

We analyze bias in historical corpora as encoded in diachronic distributional semantic models by focusing on two specific forms of bias, namely a political (i.e., anti-communism) and racist (i.e., antisemitism) one. For this, we use a new…

Computation and Language · Computer Science 2021-08-16 Tobias Walter , Celina Kirschner , Steffen Eger , Goran Glavaš , Anne Lauscher , Simone Paolo Ponzetto

Parliamentary transcripts provide a valuable resource to understand the reality and know about the most important facts that occur over time in our societies. Furthermore, the political debates captured in these transcripts facilitate…

Current research into spoken language translation (SLT),or speech-to-text translation, is often hampered by the lack of specific data resources for this task, as currently available SLT datasets are restricted to a limited set of language…

Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread…

Computation and Language · Computer Science 2026-04-14 Lena S. Oberkircher , Jesujoba O. Alabi , Dietrich Klakow , Jürgen Trouvain

We introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel. To the best of our knowledge, this is the first single…

Computation and Language · Computer Science 2022-05-25 Debjoy Saha , Shravan Nayak , Timo Baumann

This paper examines the current state-of-the-art of German text simplification, focusing on parallel and monolingual German corpora. It reviews neural language models for simplifying German texts and assesses their suitability for legal…

Computation and Language · Computer Science 2023-12-18 Thorben Schomacker , Michael Gille , Jörg von der Hülls , Marina Tropmann-Frick

This paper presents the "Leipzig Corpus Miner", a technical infrastructure for supporting qualitative and quantitative content analysis. The infrastructure aims at the integration of 'close reading' procedures on individual documents with…

Computation and Language · Computer Science 2017-07-12 Andreas Niekler , Gregor Wiedemann , Gerhard Heyer

The growing body of political texts opens up new opportunities for rich insights into political dynamics and ideologies but also increases the workload for manual analysis. Automated speaker attribution, which detects who said what to whom…

Computation and Language · Computer Science 2024-03-14 Tobias Bornheim , Niklas Grieger , Patrick Gustav Blaneck , Stephan Bialonski

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to talk shows and…

Computation and Language · Computer Science 2025-09-25 Samuel Stucki , Mark Cieliebak , Jan Deriu

Swiss German is a dialect continuum whose natively acquired dialects significantly differ from the formal variety of the language. These dialects are mostly used for verbal communication and do not have standard orthography. This has led to…

Computation and Language · Computer Science 2021-03-23 Pelin Dogan-Schönberger , Julian Mäder , Thomas Hofmann

This study investigates political discourse in the German parliament, the Bundestag, by analyzing approximately 28,000 parliamentary speeches from the last five years. Two machine learning models for topic and sentiment classification were…

Computation and Language · Computer Science 2025-08-06 Lukas Pätz , Moritz Beyer , Jannik Späth , Lasse Bohlen , Patrick Zschech , Mathias Kraus , Julian Rosenberger
‹ Prev 1 2 3 10 Next ›