English
Related papers

Related papers: German Parliamentary Corpus (GerParCor)

200 papers

The increasing availability of audio data on the internet lead to a multitude of datasets for development and training of text to speech applications, based on neural networks. Highly differing quality of voice, low sampling rates, lack of…

Sound · Computer Science 2021-06-14 Pascal Puchtler , Johannes Wirth , René Peinl

This paper presents SwissCrawl, the largest Swiss German text corpus to date. Composed of more than half a million sentences, it was generated using a customized web scraping tool that could be applied to other low-resource languages as…

Computation and Language · Computer Science 2020-06-17 Lucy Linder , Michael Jungo , Jean Hennebert , Claudiu Musat , Andreas Fischer

Neural Machine Translation with its significant results, still has a great problem: lack or absence of parallel corpus for many languages. This article suggests a method for generating considerable amount of parallel corpus for any language…

Computation and Language · Computer Science 2018-04-12 Farshad Jafari

This study presents German FinBERT, a novel pre-trained German language model tailored for financial textual data. The model is trained through a comprehensive pre-training process, leveraging a substantial corpus comprising financial…

Computation and Language · Computer Science 2023-11-16 Moritz Scherrmann

Migration has been a core topic in German political debate, from the postwar displacement of millions of expellees to labor migration and recent refugee movements. Studying political speech across such wide-ranging phenomena in depth has…

Computation and Language · Computer Science 2026-04-06 Aida Kostikova , Ole Pütz , Steffen Eger , Olga Sabelfeld , Benjamin Paassen

Evaluating large language models (LLMs) for legal reasoning requires workflows that span task design, expert annotation, model execution, and metric-based evaluation. In practice, these steps are split across platforms and scripts, limiting…

Computation and Language · Computer Science 2026-05-28 Sebastian Nagl , Matthias Grabmair

We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present in any other LLM…

Computation and Language · Computer Science 2026-04-02 Jesse van Oort , Frank Brinkkemper , Erik de Graaf , Bram Vanroy , Saskia Lensink

Automatic text simplification systems help to reduce textual information barriers on the internet. However, for languages other than English, only few parallel data to train these systems exists. We propose a two-step approach to overcome…

Computation and Language · Computer Science 2023-11-08 Miriam Anschütz , Joshua Oehms , Thomas Wimmer , Bartłomiej Jezierski , Georg Groh

Large, diachronic datasets of political discourse are hard to come across, especially for resource-lean languages such as Greek. In this paper, we introduce a curated dataset of the Greek Parliament Proceedings that extends chronologically…

Computation and Language · Computer Science 2022-10-25 Konstantina Dritsa , Kaiti Thoma , John Pavlopoulos , Panos Louridas

The paper presents a new training dataset of sentences in 7 languages, manually annotated for sentiment, which are used in a series of experiments focused on training a robust sentiment identifier for parliamentary proceedings. The paper…

Computation and Language · Computer Science 2024-03-21 Michal Mochtak , Peter Rupnik , Nikola Ljubešić

This paper accompanies the release of Opusparcus, a new paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish. The corpus consists of paraphrases, that is, pairs of sentences in the same…

Computation and Language · Computer Science 2018-09-18 Mathias Creutz

This article introduces to the interactive Leipzig Corpus Miner (iLCM) - a newly released, open-source software to perform automatic content analysis. Since the iLCM is based on the R-programming language, its generic text mining procedures…

Computation and Language · Computer Science 2021-10-07 Christian Kahmann , Andreas Niekler , Gregor Wiedemann

This paper describes an interdisciplinary approach which brings together the fields of corpus linguistics and translation studies. It presents ongoing work on the creation of a corpus resource in which translation shifts are explicitly…

Computation and Language · Computer Science 2007-05-23 Lea Cyrus

This article explores the requirements for corpus compilation within the GiesKaNe project (University of Giessen and Kassel, Syntactic Basic Structures of New High German). The project is defined by three central characteristics: it is a…

Computation and Language · Computer Science 2025-02-10 Volker Emmrich

In a representative democracy, some decide in the name of the rest, and these elected officials are commonly gathered in public assemblies, such as parliaments, where they discuss policies, legislate, and vote on fundamental initiatives. A…

Computation and Language · Computer Science 2022-07-01 Paulo Almeida , Manuel Marques-Pita , Joana Gonçalves-Sá

We present a new, unique and freely available parallel corpus containing European Union (EU) documents of mostly legal nature. It is available in all 20 official EUanguages, with additional documents being available in the languages of the…

Computation and Language · Computer Science 2007-05-23 Ralf Steinberger , Bruno Pouliquen , Anna Widiger , Camelia Ignat , Tomaz Erjavec , Dan Tufis , Daniel Varga

We present STT4SG-350 (Speech-to-Text for Swiss German), a corpus of Swiss German speech, annotated with Standard German text at the sentence level. The data is collected using a web app in which the speakers are shown Standard German…

We survey clinical document corpora, with focus on German textual data. Due to rigid data privacy legislation in Germany these resources, with only few exceptions, are stored in safe clinical data spaces and locked against clinic-external…

Computation and Language · Computer Science 2025-02-20 Udo Hahn

The electoral programs of six German parties issued before the parliamentary elections of 2021 are analyzed using state-of-the-art computational tools for quantitative narrative, topic and sentiment analysis. We compare different methods…

Computation and Language · Computer Science 2021-09-28 Arthur M. Jacobs , Annette Kinder

Subjects change frequently in moderated debates with several participants, such as in parliamentary sessions, electoral debates, and trials. Partitioning a debate into blocks with the same subject is essential for understanding. Often a…