中文
相关论文

相关论文: Non-verbal information in spontaneous speech -- to…

200 篇论文

In conversational speech, the acoustic signal provides cues that help listeners disambiguate difficult parses. For automatically parsing spoken utterances, we introduce a model that integrates transcribed text and acoustic-prosodic features…

计算与语言 · 计算机科学 2018-04-17 Trang Tran , Shubham Toshniwal , Mohit Bansal , Kevin Gimpel , Karen Livescu , Mari Ostendorf

In spoken communication, information is transmitted not only via words, but also through a rich array of non-verbal signals, including prosody--the non-segmental auditory features of speech. Do these different communication channels carry…

A crucial step in processing speech audio data for information extraction, topic detection, or browsing/playback is to segment the input into sentence and topic units. Speech segmentation is challenging, since the cues typically present for…

计算与语言 · 计算机科学 2022-02-28 E. Shriberg , A. Stolcke , D. Hakkani-Tur , G. Tur

The prosodic aspects of speech signals produced by current text-to-speech systems are typically averaged over training material, and as such lack the variety and liveliness found in natural speech. To avoid monotony and averaged prosody…

计算与语言 · 计算机科学 2019-06-05 Vincent Wan , Chun-an Chan , Tom Kenter , Jakub Vit , Rob Clark

Prosody -- the melody of speech -- conveys critical information often not captured by the words or text of a message. In this paper, we propose an information-theoretic approach to quantify how much information is expressed by prosody alone…

计算与语言 · 计算机科学 2025-12-19 Aditya Yadavalli , Tiago Pimentel , Tamar I Regev , Ethan Wilcox , Alex Warstadt

Identifying whether an utterance is a statement, question, greeting, and so forth is integral to effective automatic understanding of natural dialog. Little is known, however, about how such dialog acts (DAs) can be automatically classified…

计算与语言 · 计算机科学 2007-05-23 E. Shriberg , R. Bates , A. Stolcke , P. Taylor , D. Jurafsky , K. Ries , N. Coccaro , R. Martin , M. Meteer , C. Van Ess-Dykema

We address the problem of inferring a speaker's level of certainty based on prosodic information in the speech signal, which has application in speech-based dialogue systems. We show that using phrase-level prosodic features centered around…

计算与语言 · 计算机科学 2011-03-11 Heather Pon-Barry , Stuart M. Shieber

Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody varies at a slower rate compared with other content in the…

音频与语音处理 · 电气工程与系统科学 2021-02-15 Zack Hodari , Alexis Moinet , Sri Karlapati , Jaime Lorenzo-Trueba , Thomas Merritt , Arnaud Joly , Ammar Abbas , Penny Karanasou , Thomas Drugman

Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content and speaker identity…

声音 · 计算机科学 2023-09-27 Leyuan Qu , Taihao Li , Cornelius Weber , Theresa Pekarek-Rosin , Fuji Ren , Stefan Wermter

Spoken language understanding research to date has generally carried a heavy text perspective. Most datasets are derived from text, which is then subsequently synthesized into speech, and most models typically rely on automatic…

计算与语言 · 计算机科学 2025-02-11 Jie Chi , Maureen de Seyssel , Natalie Schluter

People exploit the predictability of lexical structures during text comprehension. Though predictable structure is also present in speech, the degree to which prosody, e.g. intonation, tempo, and loudness, contributes to such structure…

计算与语言 · 计算机科学 2025-06-04 Sarenne Wallbridge , Christoph Minixhofer , Catherine Lai , Peter Bell

In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However, when learning to control prosody in text-to-speech voices,…

音频与语音处理 · 电气工程与系统科学 2020-11-04 Zack Hodari , Catherine Lai , Simon King

Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used in downstream tasks…

音频与语音处理 · 电气工程与系统科学 2025-08-14 Eray Eren , Qingju Liu , Hyeongwoo Kim , Pablo Garrido , Abeer Alwan

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversion. The contrastive…

音频与语音处理 · 电气工程与系统科学 2024-09-06 Yuying Xie , Michael Kuhlmann , Frederik Rautenberg , Zheng-Hua Tan , Reinhold Haeb-Umbach

The differences in written text and conversational speech are substantial; previous parsers trained on treebanked text have given very poor results on spontaneous speech. For spoken language, the mismatch in style also extends to prosodic…

计算与语言 · 计算机科学 2020-10-12 Trang Tran , Jiahong Yuan , Yang Liu , Mari Ostendorf

Prosody -- the suprasegmental component of speech, including pitch, loudness, and tempo -- carries critical aspects of meaning. However, the relationship between the information conveyed by prosody vs. by the words themselves remains poorly…

计算与语言 · 计算机科学 2023-11-30 Lukas Wolf , Tiago Pimentel , Evelina Fedorenko , Ryan Cotterell , Alex Warstadt , Ethan Wilcox , Tamar Regev

We propose a method for learning de-identified prosody representations from raw audio using a contrastive self-supervised signal. Whereas prior work has relied on conditioning models on bottlenecks, we introduce a set of inductive biases…

计算与语言 · 计算机科学 2021-07-20 Jack Weston , Raphael Lenain , Udeepa Meepegama , Emil Fristed

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to learn prosody…

音频与语音处理 · 电气工程与系统科学 2023-06-21 Jakub Swiatkowski , Duo Wang , Mikolaj Babianski , Patrick Lumban Tobing , Ravichander Vipperla , Vincent Pollet

Turn-taking is a fundamental aspect of human communication and can be described as the ability to take turns, project upcoming turn shifts, and supply backchannels at appropriate locations throughout a conversation. In this work, we…

音频与语音处理 · 电气工程与系统科学 2022-09-13 Erik Ekstedt , Gabriel Skantze

Dialogue disentanglement aims to detach the chronologically ordered utterances into several independent sessions. Conversation utterances are essentially organized and described by the underlying discourse, and thus dialogue disentanglement…

计算与语言 · 计算机科学 2023-06-13 Bobo Li , Hao Fei , Fei Li , Shengqiong Wu , Lizi Liao , Yinwei Wei , Tat-Seng Chua , Donghong Ji
‹ 上一页 1 2 3 10 下一页 ›