English
Related papers

Related papers: German Parliamentary Corpus (GerParCor)

200 papers

Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets…

Traditionally, Text Simplification is treated as a monolingual translation task where sentences between source texts and their simplified counterparts are aligned for training. However, especially for longer input documents, summarizing the…

Computation and Language · Computer Science 2022-07-29 Dennis Aumiller , Michael Gertz

We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article…

Computation and Language · Computer Science 2025-05-01 Michelle Wastl , Jannis Vamvas , Selena Calleri , Rico Sennrich

A major challenge of research on non-English machine reading for question answering (QA) is the lack of annotated datasets. In this paper, we present GermanQuAD, a dataset of 13,722 extractive question/answer pairs. To improve the…

Computation and Language · Computer Science 2021-04-27 Timo Möller , Julian Risch , Malte Pietsch

Sociality borne by language, as is the predominant digital trace on text-based social media platforms, harbours the raw material for exploring a multitude of social phenomena. Distinctively, the messaging service Telegram provides…

Social and Information Networks · Computer Science 2026-04-28 Mathias Angermaier , Elisabeth Hoeldrich , Jana Lasser , Joao Pinheiro Neto

Quantitative Discourse Analysis has seen growing adoption with the rise of Large Language Models and computational tools. However, reliance on black box software such as MAXQDA and NVivo risks undermining methodological transparency and…

Computation and Language · Computer Science 2025-08-27 Thomas Compton

This paper targets the automated extraction of components of argumentative information and their relations from natural language text. Moreover, we address a current lack of systems to provide complete argumentative structure from arbitrary…

Computation and Language · Computer Science 2020-09-29 Mirko Lenz , Premtim Sahitaj , Sean Kallenberg , Christopher Coors , Lorik Dumani , Ralf Schenkel , Ralph Bergmann

LLMs are ubiquitous in modern NLP, and while their applicability extends to texts produced for democratic activities such as online deliberations or large-scale citizen consultations, ethical questions have been raised for their usage as…

Computation and Language · Computer Science 2026-04-21 Pierre-Antoine Lequeu , Léo Labat , Laurène Cave , Gaël Lejeune , François Yvon , Benjamin Piwowarski

Research in Computational Linguistics is dependent on text corpora for training and testing new tools and methodologies. While there exists a plethora of annotated linguistic information, these corpora are often not interoperable without…

Computation and Language · Computer Science 2020-11-03 Timo Lek , Anna de Groot , Tobias Kuhn , Roser Morante

This paper introduces GigaST, a large-scale pseudo speech translation (ST) corpus. We create the corpus by translating the text in GigaSpeech, an English ASR corpus, into German and Chinese. The training set is translated by a strong…

Computation and Language · Computer Science 2023-06-07 Rong Ye , Chengqi Zhao , Tom Ko , Chutong Meng , Tao Wang , Mingxuan Wang , Jun Cao

Text similarity detection aims at measuring the degree of similarity between a pair of texts. Corpora available for text similarity detection are designed to evaluate the algorithms to assess the paraphrase level among documents. In this…

Information Retrieval · Computer Science 2017-03-14 Juan-Manuel Torres-Moreno , Gerardo Sierra , Peter Peinl

We present the Multilingual TEDx corpus, built to support speech recognition (ASR) and speech translation (ST) research across many non-English source languages. The corpus is a collection of audio recordings from TEDx talks in 8 source…

Computation and Language · Computer Science 2021-06-16 Elizabeth Salesky , Matthew Wiesner , Jacob Bremerman , Roldano Cattoni , Matteo Negri , Marco Turchi , Douglas W. Oard , Matt Post

Extracting coherent and human-understandable themes from large collections of unstructured historical newspaper archives presents significant challenges due to topic evolution, Optical Character Recognition (OCR) noise, and the sheer volume…

Computation and Language · Computer Science 2025-12-15 Keerthana Murugaraj , Salima Lamsiyah , Marten During , Martin Theobald

Lectures translation is a case of spoken language translation and there is a lack of publicly available parallel corpora for this purpose. To address this, we examine a language independent framework for parallel corpus mining which is a…

Computation and Language · Computer Science 2020-01-15 Haiyue Song , Raj Dabre , Atsushi Fujita , Sadao Kurohashi

Being able to understand information is a key factor for a self-determined life and society. It is also very important for participating in democratic processes. The study of automatic text simplification is often limited by the…

Computation and Language · Computer Science 2026-03-17 Stefan Bott , Verena Riegler , Horacio Saggion , Almudena Rascón Alcaina , Nouran Khallaf

Recent advances in the area of legal information systems have led to a variety of applications that promise support in processing and accessing legal documents. Unfortunately, these applications have various limitations, e.g., regarding…

Digital Libraries · Computer Science 2020-05-28 Malte Ostendorff , Till Blume , Saskia Ostendorff

We introduce an advanced information extraction pipeline to automatically process very large collections of unstructured textual data for the purpose of investigative journalism. The pipeline serves as a new input processor for the upcoming…

Computation and Language · Computer Science 2018-09-17 Gregor Wiedemann , Seid Muhie Yimam , Chris Biemann

The study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages. Tackling these…

Computation and Language · Computer Science 2025-03-25 Farhan Farsi , Parnian Fazel , Sepand Haghighi , Sadra Sabouri , Farzaneh Goshtasb , Nadia Hajipour , Ehsaneddin Asgari , Hossein Sameti

Multilinguality is gradually becoming ubiquitous in the sense that more and more researchers have successfully shown that using additional languages help improve the results in many Natural Language Processing tasks. Multilingual Multiway…

Computation and Language · Computer Science 2019-02-15 Raj Dabre , Sadao Kurohashi

Public knowledge of what is said in parliament is a tenet of democracy, and a critical resource for political science research. In Australia, following the British tradition, the written record of what is said in parliament is known as…

Digital Libraries · Computer Science 2023-09-25 Lindsay Katz , Rohan Alexander
‹ Prev 1 4 5 6 7 8 10 Next ›