English
Related papers

Related papers: Many uses, many annotations for large speech corpo…

200 papers

Discourse information is difficult to represent and annotate. Among the major frameworks for annotating discourse information, RST, PDTB and SDRT are widely discussed and used, each having its own theoretical foundation and focus. Corpora…

Computation and Language · Computer Science 2022-04-19 Yingxue Fu

Recent trends in natural language processing research and annotation tasks affirm a paradigm shift from the traditional reliance on a single ground truth to a focus on individual perspectives, particularly in subjective tasks. In scenarios…

Computation and Language · Computer Science 2024-04-18 Olufunke O. Sarumi , Béla Neuendorf , Joan Plepi , Lucie Flek , Jörg Schlötterer , Charles Welch

Annotated speech corpora are databases consisting of signal data along with time-aligned symbolic `transcriptions'. Such databases are typically multidimensional, heterogeneous and dynamic. These properties present a number of tough…

Computation and Language · Computer Science 2007-05-23 Steve Cassidy , Steven Bird

Transitioning between topics is a natural component of human-human dialog. Although topic transition has been studied in dialogue for decades, only a handful of corpora based studies have been performed to investigate the subtleties of…

Computation and Language · Computer Science 2022-07-21 Mayank Soni , Brendan Spillane , Emer Gilmartin , Christian Saam , Benjamin R. Cowan , Vincent Wade

Discourse-annotated corpora are an important resource for the community, but they are often annotated according to different frameworks. This makes comparison of the annotations difficult, thereby also preventing researchers from searching…

Computation and Language · Computer Science 2018-03-16 Vera Demberg , Fatemeh Torabi Asr , Merel Scholman

Automated fact-checking based on machine learning is a promising approach to identify false information distributed on the web. In order to achieve satisfactory performance, machine learning methods require a large corpus with reliable…

Computation and Language · Computer Science 2019-11-05 Andreas Hanselowski , Christian Stab , Claudia Schulz , Zile Li , Iryna Gurevych

This paper describes the Spot the Difference Corpus which contains 54 interactions between pairs of subjects interacting to find differences in two very similar scenes. The setup used, the participants' metadata and details about collection…

Computation and Language · Computer Science 2018-05-15 José Lopes , Nils Hemmingsson , Oliver Åstrand

Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods relying on speech prompts (reference speech) for voice…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-13 Yichong Leng , Zhifang Guo , Kai Shen , Xu Tan , Zeqian Ju , Yanqing Liu , Yufei Liu , Dongchao Yang , Leying Zhang , Kaitao Song , Lei He , Xiang-Yang Li , Sheng Zhao , Tao Qin , Jiang Bian

This paper explores contexts associated with errors in transcrip-tion of spontaneous speech, shedding light on human perceptionof disfluencies and other conversational speech phenomena. Anew version of the Switchboard corpus is provided…

Computation and Language · Computer Science 2019-04-10 Vicky Zayats , Trang Tran , Richard Wright , Courtney Mansfield , Mari Ostendorf

Much linguistic research relies on annotated datasets of features extracted from text corpora, but the rapid quantitative growth of these corpora has created practical difficulties for linguists to manually annotate large data samples. In…

Computation and Language · Computer Science 2025-04-11 Cameron Morin , Matti Marttinen Larsson

Disentangling conversations mixed together in a single stream of messages is a difficult task, made harder by the lack of large manually annotated datasets. We created a new dataset of 77,563 messages manually annotated with reply-structure…

Properly annotated multimedia content is crucial for supporting advances in many Information Retrieval applications. It enables, for instance, the development of automatic tools for the annotation of large and diverse multimedia…

Information Retrieval · Computer Science 2018-11-28 Xavier Favory , Eduardo Fonseca , Frederic Font , Xavier Serra

This paper presents the challenges in creating and managing large parallel corpora of 12 major Indian languages (which is soon to be extended to 23 languages) as part of a major consortium project funded by the Department of Information…

Computation and Language · Computer Science 2021-12-06 Ritesh Kumar , Shiv Bhusan Kaushik , Pinkey Nainwani , Girish Nath Jha

Research in Computational Linguistics is dependent on text corpora for training and testing new tools and methodologies. While there exists a plethora of annotated linguistic information, these corpora are often not interoperable without…

Computation and Language · Computer Science 2020-11-03 Timo Lek , Anna de Groot , Tobias Kuhn , Roser Morante

Existing discourse corpora are annotated based on different frameworks, which show significant dissimilarities in definitions of arguments and relations and structural constraints. Despite surface differences, these frameworks share basic…

Computation and Language · Computer Science 2024-04-09 Yingxue Fu

Data-to-text (D2T) and text-to-data (T2D) are dual tasks that convert structured data, such as graphs or tables into fluent text, and vice versa. These tasks are usually handled separately and use corpora extracted from a single source.…

Machine Learning · Computer Science 2023-02-23 Song Duong , Alberto Lumbreras , Mike Gartrell , Patrick Gallinari

This paper elaborates on the notion of uncertainty in the context of annotation in large text corpora, specifically focusing on (but not limited to) historical languages. Such uncertainty might be due to inherent properties of the language,…

Computation and Language · Computer Science 2021-05-31 Marie-Luis Merten , Marcel Wever , Michaela Geierhos , Doris Tophinke , Eyke Hüllermeier

This paper has two goals. First, we present the turn-taking annotation layers created for 95 minutes of conversational speech of the Graz Corpus of Read and Spontaneous Speech (GRASS), available to the scientific community. Second, we…

Computation and Language · Computer Science 2025-04-15 Anneliese Kelterer , Barbara Schuppler

This paper describes an interdisciplinary approach which brings together the fields of corpus linguistics and translation studies. It presents ongoing work on the creation of a corpus resource in which translation shifts are explicitly…

Computation and Language · Computer Science 2007-05-23 Lea Cyrus

This article presents an analysis of the influence of context information on dialog act recognition. We performed experiments on the widely explored Switchboard corpus, as well as on data annotated according to the recent ISO 24617-2…

Computation and Language · Computer Science 2017-01-10 Eugénio Ribeiro , Ricardo Ribeiro , David Martins de Matos
‹ Prev 1 2 3 10 Next ›