中文
相关论文

相关论文: German Parliamentary Corpus (GerParCor)

200 篇论文

Over the past years, interest in discourse analysis and discourse parsing has steadily grown, and many discourse-annotated corpora and, as a result, discourse parsers have been built. In this paper, we present a discourse-annotated corpus…

计算与语言 · 计算机科学 2021-06-29 Sara Shahmohammadi , Hadi Veisi , Ali Darzi

This paper reports about experiments with GermaNet as a resource within domain specific document analysis. The main question to be answered is: How is the coverage of GermaNet in a specific domain? We report about results of a field test of…

人工智能 · 计算机科学 2007-05-23 Manuela Kunze , Dietmar Roesner

Sarcasm is a complex form of figurative language in which the intended meaning contradicts the literal one. Its prevalence in social media and popular culture poses persistent challenges for natural language understanding, sentiment…

计算与语言 · 计算机科学 2026-03-05 Aaron Scott , Maike Züfle , Jan Niehues

The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce German4All, the first large-scale German dataset of aligned…

Lately, pre-trained language models advanced the field of natural language processing (NLP). The introduction of Bidirectional Encoders for Transformers (BERT) and its optimized version RoBERTa have had significant impact and increased the…

计算与语言 · 计算机科学 2025-06-13 Raphael Scheible , Fabian Thomczyk , Patric Tippmann , Victor Jaravine , Martin Boeker

In this paper, we propose the scheme for annotating large-scale multi-party chat dialogues for discourse parsing and machine comprehension. The main goal of this project is to help understand multi-party dialogues. Our dataset is based on…

计算与语言 · 计算机科学 2019-11-12 Jiaqi Li , Ming Liu , Bing Qin , Zihao Zheng , Ting Liu

Despite the success of attention-based neural models for natural language generation and classification tasks, they are unable to capture the discourse structure of larger documents. We hypothesize that explicit discourse representations…

计算与语言 · 计算机科学 2019-11-19 Fajri Koto , Jey Han Lau , Timothy Baldwin

As part of its digitization initiative, the German Central Bank (Deutsche Bundesbank) wants to examine the extent to which natural Language Processing (NLP) can be used to make independent decisions upon the eligibility criteria of…

计算与语言 · 计算机科学 2023-02-10 Christian Hänig , Markus Schlösser , Serhii Hamotskyi , Gent Zambaku , Janek Blankenburg

This work introduces a benchmark assessing the performance of clustering German text embeddings in different domains. This benchmark is driven by the increasing use of clustering neural text embeddings in tasks that require the grouping of…

计算与语言 · 计算机科学 2024-01-08 Silvan Wehrli , Bert Arnrich , Christopher Irrgang

With the increasing popularity of mobile devices and the wide adoption of mobile Apps, an increasing concern of privacy issues is raised. Privacy policy is identified as a proper medium to indicate the legal terms, such as GDPR, and to bind…

计算机与社会 · 计算机科学 2020-05-15 Shuang Liu , Renjie Guo , Baiyang Zhao , Tao Chen , Meishan Zhang

Narratives are key interpretative devices by which humans make sense of political reality. As the significance of narratives for understanding current societal issues such as polarization and misinformation becomes increasingly evident,…

计算与语言 · 计算机科学 2025-11-10 Armin Pournaki , Tom Willaert

We focus on multi-turn response selection in a retrieval-based dialog system. In this paper, we utilize the powerful pre-trained language model Bi-directional Encoder Representations from Transformer (BERT) for a multi-turn dialog system…

计算与语言 · 计算机科学 2020-07-28 Taesun Whang , Dongyub Lee , Chanhee Lee , Kisu Yang , Dongsuk Oh , HeuiSeok Lim

We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information…

计算与语言 · 计算机科学 2025-09-09 Jinrui Yang , Timothy Baldwin , Trevor Cohn

Recent progress in natural language processing has been impressive in many different areas with transformer-based approaches setting new benchmarks for a wide range of applications. This development has also lowered the barriers for people…

计算与语言 · 计算机科学 2022-04-07 Miriam Schirmer , Udo Kruschwitz , Gregor Donabauer

We describe a gold standard corpus of protest events that comprise of various local and international sources from various countries in English. The corpus contains document, sentence, and token level annotations. This corpus facilitates…

计算与语言 · 计算机科学 2020-08-04 Ali Hürriyetoğlu , Erdem Yörük , Deniz Yüret , Osman Mutlu , Çağrı Yoltar , Fırat Duruşan , Burak Gürel

There is a significant disconnect between linguistic theory and modern NLP practice, which relies heavily on inscrutable black-box architectures. DisCoCirc is a newly proposed model for meaning that aims to bridge this divide, by providing…

计算与语言 · 计算机科学 2023-11-30 Jonathon Liu , Razin A. Shaikh , Benjamin Rodatz , Richie Yeung , Bob Coecke

This study is devoted to the automatic generation of German drama texts. We suggest an approach consisting of two key steps: fine-tuning a GPT-2 model (the outline model) to generate outlines of scenes based on keywords and fine-tuning a…

计算与语言 · 计算机科学 2023-01-11 Mariam Bangura , Kristina Barabashova , Anna Karnysheva , Sarah Semczuk , Yifan Wang

At a time when the quantity of - more or less freely - available data is increasing significantly, thanks to digital corpora, editions or libraries, the development of data mining tools or deep learning methods allows researchers to build a…

计算机视觉与模式识别 · 计算机科学 2019-04-29 Jean-Baptiste Camps , Gilles Guilhem Couffignal

LexNLP is an open source Python package focused on natural language processing and machine learning for legal and regulatory text. The package includes functionality to (i) segment documents, (ii) identify key text such as titles and…

计算与语言 · 计算机科学 2018-06-12 Michael J Bommarito , Daniel Martin Katz , Eric M Detterman

Large language models (LLMs) are increasingly shaping citizens' information ecosystems. Products incorporating LLMs, such as chatbots and AI Companions, are now widely used for decision support and information retrieval, including in…

计算机与社会 · 计算机科学 2025-12-02 Jan Batzner , Volker Stocker , Stefan Schmid , Gjergji Kasneci