中文
相关论文

相关论文: What Do Prosody and Text Convey? Characterizing Ho…

200 篇论文

While there has been significant progress towards modelling coherence in written discourse, the work in modelling spoken discourse coherence has been quite limited. Unlike the coherence in text, coherence in spoken discourse is also…

计算与语言 · 计算机科学 2021-01-05 Rajaswa Patil , Yaman Kumar Singla , Rajiv Ratn Shah , Mika Hama , Roger Zimmermann

Psycholinguistic studies of human word processing and lexical access provide ample evidence of the preferred nature of word-initial versus word-final segments, e.g., in terms of attention paid by listeners (greater) or the likelihood of…

计算与语言 · 计算机科学 2021-02-04 Tiago Pimentel , Ryan Cotterell , Brian Roark

A sharp tension exists about the nature of human language between two opposite parties: those who believe that statistical surface distributions, in particular using measures like surprisal, provide a better understanding of language…

计算与语言 · 计算机科学 2023-02-20 Matteo Greco , Andrea Cometa , Fiorenzo Artoni , Robert Frank , Andrea Moro

Self-supervised representation learning for speech often involves a quantization step that transforms the acoustic input into discrete units. However, it remains unclear how to characterize the relationship between these discrete units and…

计算与语言 · 计算机科学 2023-06-06 Badr M. Abdullah , Mohammed Maqsood Shaik , Bernd Möbius , Dietrich Klakow

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal

Many popular form factors of digital assistants---such as Amazon Echo, Apple Homepod, or Google Home---enable the user to hold a conversation with these systems based only on the speech modality. The lack of a screen presents unique…

计算与语言 · 计算机科学 2019-10-03 Aleksandr Chuklin , Aliaksei Severyn , Johanne Trippas , Enrique Alfonseca , Hanna Silen , Damiano Spina

The way that humans encode their emotion into speech signals is complex. For instance, an angry man may increase his pitch and speaking rate, and use impolite words. In this paper, we present a preliminary study on various emotional factors…

声音 · 计算机科学 2021-11-25 Haoran Sun , Lantian Li , Thomas Fang Zheng , Dong Wang

Collecting everyday speech data for prosodic analysis is challenging due to the confounding of prosody and semantics, privacy constraints, and participant compliance. We introduce and empirically evaluate a content-controlled, privacy-first…

人机交互 · 计算机科学 2026-03-19 Timo K. Koch , Florian Bemmann , Ramona Schoedel , Markus Buehner , Clemens Stachl

We present a framework to identify whether a public speaker's body movements are meaningful or non-meaningful ("Mannerisms") in the context of their speeches. In a dataset of 84 public speaking videos from 28 individuals, we extract 314…

人机交互 · 计算机科学 2017-07-18 Md Iftekhar Tanveer , RuJie Zhao , Mohammed Hoque

In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, discarding prosodic…

人机交互 · 计算机科学 2026-03-11 SangYeop Jeong , Yeongseo Na , Seung Gyu Jeong , Jin-Woo Jeong , Seong-Eun Kim

We examine how well people learn when information is noisily relayed from person to person; and we study how communication platforms can improve learning without censoring or fact-checking messages. We analyze learning as a function of…

物理与社会 · 物理学 2020-06-30 Matthew O. Jackson , Suraj Malladi , David McAdams

Authors of posts in social media communicate their emotions and what causes them with text and images. While there is work on emotion and stimulus detection for each modality separately, it is yet unknown if the modalities contain…

计算机视觉与模式识别 · 计算机科学 2022-05-17 Anna Khlyzova , Carina Silberer , Roman Klinger

We suggest an information-theoretic approach for measuring stylistic coordination in dialogues. The proposed measure has a simple predictive interpretation and can account for various confounding factors through proper conditioning. We…

计算与语言 · 计算机科学 2015-08-28 Shuyang Gao , Greg Ver Steeg , Aram Galstyan

Consonance is related to the perception of pleasantness arising from a combination of sounds and has been approached quantitatively using mathematical relations, physics, information theory, and psychoacoustics. Tonal consonance is present…

声音 · 计算机科学 2017-04-25 Jorge Useche , Rafael Hurtado

The study of inter-human communication requires a more complex framework than Shannon's (1948) mathematical theory of communication because "information" is defined in the latter case as meaningless uncertainty. Assuming that meaning cannot…

数字图书馆 · 计算机科学 2013-03-15 Loet Leydesdorff , Inga A. Ivanova

We reconsider the persistence of information under the dynamics of the logistic map in order to discuss communication through a nonlinear channel where the sender can set the initial state of the system with finite resolution, and the…

混沌动力学 · 物理学 2009-11-10 Richard Metzler , Yaneer Bar-Yam , Mehran Kardar

We address the problem of human-in-the-loop control for generating prosody in the context of text-to-speech synthesis. Controlling prosody is challenging because existing generative models lack an efficient interface through which users can…

音频与语音处理 · 电气工程与系统科学 2024-04-17 Dan Andrei Iliescu , Devang Savita Ram Mohan , Tian Huey Teh , Zack Hodari

Human emotional expression emerges through coordinated vocal, facial, and gestural signals. While speech face alignment is well established, the broader dynamics linking emotionally expressive speech to regional facial and hand motion…

多媒体 · 计算机科学 2025-06-13 Von Ralph Dane Marquez Herbuela , Yukie Nagai

Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating…

音频与语音处理 · 电气工程与系统科学 2026-03-23 Lokesh Kumar , Nirmesh Shah , Ashishkumar P. Gudmalwar , Pankaj Wasnik

Interpreting uncertain data can be difficult, particularly if the data presentation is complex. We investigate the efficacy of different modalities for representing data and how to combine the strengths of each modality to facilitate the…

人机交互 · 计算机科学 2024-04-15 Chase Stokes , Chelsea Sanker , Bridget Cogley , Vidya Setlur