中文
相关论文

相关论文: MaiBaam Annotation Guidelines

200 篇论文

Despite the success of the Universal Dependencies (UD) project exemplified by its impressive language breadth, there is still a lack in `within-language breadth': most treebanks focus on standard languages. Even for German, the language…

计算与语言 · 计算机科学 2024-03-18 Verena Blaschke , Barbara Kovačić , Siyao Peng , Hinrich Schütze , Barbara Plank

Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word…

We present a lemmatizer/POS-tagger/dependency parser for West Frisian using a corpus of 44,714 words in 3,126 sentences that were annotated according to the guidelines of Universal Dependency version 2. POS tags were assigned to words by…

计算与语言 · 计算机科学 2021-07-29 Wilbert Heeringa , Gosse Bouma , Martha Hofman , Eduard Drenth , Jan Wijffels , Hans Van de Velde

This paper explores the difficulties of annotating transcribed spoken Dutch-Frisian code-switch utterances into Universal Dependencies. We make use of data from the FAME! corpus, which consists of transcriptions and audio data. Besides the…

计算与语言 · 计算机科学 2021-02-23 Anouck Braggaar , Rob van der Goot

The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns - including collaborative coconstructions proper, wh-question answers, and backchannels - in spoken language treebanks within the Universal…

Human annotations are an important source of information in the development of natural language understanding approaches. As under the pressure of productivity annotators can assign different labels to a given text, the quality of produced…

计算与语言 · 计算机科学 2020-10-29 Kristian Miok , Gregor Pirs , Marko Robnik-Sikonja

We propose a novel neural network model for joint part-of-speech (POS) tagging and dependency parsing. Our model extends the well-known BIST graph-based dependency parser (Kiperwasser and Goldberg, 2016) by incorporating a BiLSTM-based…

计算与语言 · 计算机科学 2019-08-30 Dat Quoc Nguyen , Karin Verspoor

We study the problem of analyzing tweets with Universal Dependencies. We extend the UD guidelines to cover special constructions in tweets that affect tokenization, part-of-speech tagging, and labeled dependencies. Using the extended…

计算与语言 · 计算机科学 2018-04-24 Yijia Liu , Yi Zhu , Wanxiang Che , Bing Qin , Nathan Schneider , Noah A. Smith

Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address this gap, we use…

计算与语言 · 计算机科学 2025-09-23 Robert Litschko , Verena Blaschke , Diana Burkhardt , Barbara Plank , Diego Frassinelli

The present study extends recent work on Universal Dependencies annotations for second-language (L2) Korean by introducing a semi-automated framework that identifies morphosyntactic constructions from XPOS sequences and aligns those…

计算与语言 · 计算机科学 2025-06-12 Hakyung Sung , Gyu-Ho Shin , Chanyoung Lee , You Kyung Sung , Boo Kyung Jung

Norwegian Twitter data poses an interesting challenge for Natural Language Processing (NLP) tasks. These texts are difficult for models trained on standardized text in one of the two Norwegian written forms (Bokm{\aa}l and Nynorsk), as they…

计算与语言 · 计算机科学 2022-10-13 Petter Mæhlum , Andre Kåsen , Samia Touileb , Jeremy Barnes

In this paper, we present a novel annotation approach to capture claims and premises of arguments and their relations in student-written persuasive peer reviews on business models in German language. We propose an annotation scheme based on…

计算与语言 · 计算机科学 2020-10-27 Thiemo Wambsganss , Christina Niklaus , Matthias Söllner , Siegfried Handschuh , Jan Marco Leimeister

The taggedPBC (Ring 2025a) contains more than 1,800 sentences of pos-tagged parallel text data from over 1,500 languages, representing 133 language families and 111 isolates. While this dwarfs previously available resources, and the POS…

计算与语言 · 计算机科学 2025-06-10 Hiram Ring

In this study, we aim to offer linguistically motivated solutions to resolve the issues of the lack of representation of null morphemes, highly productive derivational processes, and syncretic morphemes of Turkish in the BOUN Treebank…

This paper proposes a methodology for constructing such corpora of child directed speech (CDS) paired with sentential logical forms, and uses this method to create two such corpora, in English and Hebrew. The approach enforces a…

计算与语言 · 计算机科学 2024-03-18 Ida Szubert , Omri Abend , Nathan Schneider , Samuel Gibbon , Louis Mahon , Sharon Goldwater , Mark Steedman

The standard approach to incorporate linguistic information to neural machine translation systems consists in maintaining separate vocabularies for each of the annotated features to be incorporated (e.g. POS tags, dependency relation…

计算与语言 · 计算机科学 2021-02-18 Noe Casas , Jose A. R. Fonollosa , Marta R. Costa-jussà

With the development of big corpora of various periods, it becomes crucial to standardise linguistic annotation (e.g. lemmas, POS tags, morphological annotation) to increase the interoperability of the data produced, despite diachronic…

计算与语言 · 计算机科学 2020-11-24 Simon Gabay , Thibault Clérice , Jean-Baptiste Camps , Jean-Baptiste Tanguy , Matthias Gille-Levenson

The Universal Dependencies (UD) and Universal Morphology (UniMorph) projects each present schemata for annotating the morphosyntactic details of language. Each project also provides corpora of annotated text in many languages - UD at the…

计算与语言 · 计算机科学 2019-10-28 Arya D. McCarthy , Miikka Silfverberg , Ryan Cotterell , Mans Hulden , David Yarowsky

Multiword expressions (MWEs) refer to idiomatic sequences of multiple words. MWE identification, i.e., detecting MWEs in text, can play a key role in downstream tasks such as machine translation, but existing datasets for the task are…

计算与语言 · 计算机科学 2025-07-11 Yusuke Ide , Joshua Tanner , Adam Nohejl , Jacob Hoffman , Justin Vasselli , Hidetaka Kamigaito , Taro Watanabe

In this paper, we introduce a set of opinion annotations for the POM movie review dataset, composed of 1000 videos. The annotation campaign is motivated by the development of a hierarchical opinion prediction framework allowing one to…

多媒体 · 计算机科学 2021-04-30 Alexandre Garcia , Slim Essid , Florence d'Alché-Buc , Chloé Clavel
‹ 上一页 1 2 3 10 下一页 ›