中文
相关论文

相关论文: SpeakGer: A meta-data enriched speech corpus of Ge…

200 篇论文

This study investigates political discourse in the German parliament, the Bundestag, by analyzing approximately 28,000 parliamentary speeches from the last five years. Two machine learning models for topic and sentiment classification were…

计算与语言 · 计算机科学 2025-08-06 Lukas Pätz , Moritz Beyer , Jannik Späth , Lasse Bohlen , Patrick Zschech , Mathias Kraus , Julian Rosenberger

We introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel. To the best of our knowledge, this is the first single…

计算与语言 · 计算机科学 2022-05-25 Debjoy Saha , Shravan Nayak , Timo Baumann

We present ASR Bundestag, a dataset for automatic speech recognition in German, consisting of 610 hours of aligned audio-transcript pairs for supervised training as well as 1,038 hours of unlabeled audio snippets for self-supervised…

计算与语言 · 计算机科学 2023-02-14 Johannes Wirth , René Peinl

Large, diachronic datasets of political discourse are hard to come across, especially for resource-lean languages such as Greek. In this paper, we introduce a curated dataset of the Greek Parliament Proceedings that extends chronologically…

计算与语言 · 计算机科学 2022-10-25 Konstantina Dritsa , Kaiti Thoma , John Pavlopoulos , Panos Louridas

Parliamentary and legislative debate transcripts provide informative insight into elected politicians' opinions, positions, and policy preferences. They are interesting for political and social sciences as well as linguistics and natural…

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

计算与语言 · 计算机科学 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer

This paper presents a novel dataset of public broadcast interviews featuring high-ranking German politicians. The interviews were sourced from YouTube, transcribed, processed for speaker identification, and stored in a tidy and open format.…

计算与语言 · 计算机科学 2025-01-17 Lukas Birkenmaier , Laureen Sieber , Felix Bergstein

The electoral programs of six German parties issued before the parliamentary elections of 2021 are analyzed using state-of-the-art computational tools for quantitative narrative, topic and sentiment analysis. We compare different methods…

计算与语言 · 计算机科学 2021-09-28 Arthur M. Jacobs , Annette Kinder

The rise of populism concerns many political scientists and practitioners, yet the detection of its underlying language remains fragmentary. This paper aims to provide a reliable, valid, and scalable approach to measure populist stances.…

计算与语言 · 计算机科学 2025-01-30 L. Erhard , S. Hanke , U. Remer , A. Falenska , R. Heiberger

Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal…

计算与语言 · 计算机科学 2025-06-09 Stefanie Urchs , Veronika Thurner , Matthias Aßenmacher , Christian Heumann , Stephanie Thiemichen

The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying on principles of fairness in multilingual…

计算与语言 · 计算机科学 2026-03-12 Paul Lerner , François Yvon

Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets…

We present STT4SG-350 (Speech-to-Text for Swiss German), a corpus of Swiss German speech, annotated with Standard German text at the sentence level. The data is collected using a web app in which the speakers are shown Standard German…

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 418 thousand hours of…

Parliamentary debates represent a large and partly unexploited treasure trove of publicly accessible texts. In the German-speaking area, there is a certain deficit of uniformly accessible and annotated corpora covering all German-speaking…

计算与语言 · 计算机科学 2022-04-25 Giuseppe Abrami , Mevlüt Bagci , Leon Hammerla , Alexander Mehler

In this work, we present PoliCorp (https://demo-pollux.gesis.org/), a web portal designed to facilitate the search and analysis of political text corpora. PoliCorp provides researchers with access to rich textual data, enabling in-depth…

数字图书馆 · 计算机科学 2025-09-23 Nina Smirnova , Muhammad Ahsan Shahid , Philipp Mayr

Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary…

音频与语音处理 · 电气工程与系统科学 2026-03-06 Yanir Marmor , Arad Zulti , David Krongauz , Adam Gabet , Yoad Snapir , Yair Lifshitz , Eran Segal

The increasing digitization of political speech has opened the door to studying a new dimension of political behavior using text analysis. This work investigates the value of word-level statistical data from the US Congressional…

综合经济学 · 经济学 2018-09-05 Eitan Sapiro-Gheiler

In this paper, we present a transcribed corpus of the LIBE committee of the EU parliament, totalling 3.6 Million running words. The meetings of parliamentary committees of the EU are a potentially valuable source of information for…

计算与语言 · 计算机科学 2023-04-18 Hugo de Vos , Suzan Verberne

This paper introduces FT Speech, a new speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT). The corpus contains over 1,800 hours of transcribed speech by a total of 434 speakers.…

计算与语言 · 计算机科学 2020-10-29 Andreas Kirkedal , Marija Stepanović , Barbara Plank
‹ 上一页 1 2 3 10 下一页 ›