中文
相关论文

相关论文: What Do Prosody and Text Convey? Characterizing Ho…

200 篇论文

Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requires integrating semantic meaning (e.g., content),…

计算与语言 · 计算机科学 2026-03-17 Dingdong Wang , Junan Li , Jincenzi Wu , Dongchao Yang , Xueyuan Chen , Tianhua Zhang , Helen Meng

Speech emotion recognition (SER) is vital for obtaining emotional intelligence and understanding the contextual meaning of speech. Variations of consonant-vowel (CV) phonemic boundaries can enrich acoustic context with linguistic cues,…

声音 · 计算机科学 2023-07-03 Anna Ollerenshaw , Md Asif Jalal , Rosanna Milner , Thomas Hain

We consider the information content h of a scalar multiple-scattered, diffuse wave field $\psi(\vec{r})$ and the information capacity C of a communication channel that employs diffuse waves to transfer the information through a disordered…

无序系统与神经网络 · 物理学 2007-05-23 S. E. Skipetrov

Emotional state recognition through speech is being a very interesting research topic nowadays. Using subliminal information of speech, denominated as prosody, it is possible to recognize the emotional state of the person. One of the main…

计算机视觉与模式识别 · 计算机科学 2014-03-20 Inma Mohino-Herranz , Roberto Gil-Pita , Sagrario Alonso-Diaz , Manuel Rosa-Zurera

The speech signal conveys information on different time scales from short time scale or segmental, associated to phonological and phonetic information to long time scale or supra segmental, associated to syllabic and prosodic information.…

计算与语言 · 计算机科学 2016-09-16 Milos Cernak , Afsaneh Asaei , Hervé Bourlard

Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and datasets. We propose…

音频与语音处理 · 电气工程与系统科学 2024-12-03 Christoph Minixhofer , Ondřej Klejch , Peter Bell

The fundamental building block of social influence is for one person to elicit a response in another. Researchers measuring a "response" in social media typically depend either on detailed models of human behavior or on platform-specific…

社会与信息网络 · 计算机科学 2013-02-19 Greg Ver Steeg , Aram Galstyan

This study examines the prosodic characteristics associated with winning and losing in post-match tennis interviews. Additionally, this research explores the potential to classify match outcomes solely based on post-match interview…

计算与语言 · 计算机科学 2025-06-04 Sofoklis Kakouros , Haoyu Chen

Despite the importance of social science knowledge for various stakeholders, measuring its diffusion into different domains remains a challenge. This study uses a novel text-based approach to measure the idea-level diffusion of social…

社会与信息网络 · 计算机科学 2025-11-06 Yangliu Fan , Kilian Buehling , Volker Stocker

The interplay between nonlinear dynamic systems and noise has proved to be of great relevance in several application areas. In this presentation, we focus on the areas of information transmission and storage. We review some recent results…

其他凝聚态物理 · 物理学 2017-11-22 P. I. Fierens , G. A. Patterson , A. A. García , D. F. Grosz

Improving text representation has attracted much attention to achieve expressive text-to-speech (TTS). However, existing works only implicitly learn the prosody with masked token reconstruction tasks, which leads to low training efficiency…

声音 · 计算机科学 2023-05-19 Zhenhui Ye , Rongjie Huang , Yi Ren , Ziyue Jiang , Jinglin Liu , Jinzheng He , Xiang Yin , Zhou Zhao

Recent advances in Text-to-Speech (TTS) have improved quality and naturalness to near-human capabilities when considering isolated sentences. But something which is still lacking in order to achieve human-like communication is the dynamic…

计算与语言 · 计算机科学 2021-04-21 Shubhi Tyagi , Marco Nicolis , Jonas Rohnke , Thomas Drugman , Jaime Lorenzo-Trueba

Understanding how humans express and synchronize emotions across multiple communication channels particularly facial expressions and speech has significant implications for emotion recognition systems and human computer interaction.…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Von Ralph Dane Marquez Herbuela , Yukie Nagai

We propose prosody embeddings for emotional and expressive speech synthesis networks. The proposed methods introduce temporal structures in the embedding networks, thus enabling fine-grained control of the speaking style of the synthesized…

计算与语言 · 计算机科学 2019-02-19 Younggun Lee , Taesu Kim

Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world settings, however, the language-audio relationship is…

音频与语音处理 · 电气工程与系统科学 2025-10-22 Toranosuke Manabe , Yuchi Ishikawa , Hokuto Munakata , Tatsuya Komatsu

Developing tools to automatically detect check-worthy claims in political debates and speeches can greatly help moderators of debates, journalists, and fact-checkers. While previous work on this problem has focused exclusively on the text…

计算与语言 · 计算机科学 2024-01-19 Petar Ivanov , Ivan Koychev , Momchil Hardalov , Preslav Nakov

Identifying and communicating relationships between causes and effects is important for understanding our world, but is affected by language structure, cognitive and emotional biases, and the properties of the communication medium. Despite…

This paper introduces a novel voice conversion (VC) model, guided by text instructions such as "articulate slowly with a deep tone" or "speak in a cheerful boyish voice". Unlike traditional methods that rely on reference utterances to…

音频与语音处理 · 电气工程与系统科学 2024-01-17 Chun-Yi Kuan , Chen An Li , Tsu-Yuan Hsu , Tse-Yang Lin , Ho-Lam Chung , Kai-Wei Chang , Shuo-yiin Chang , Hung-yi Lee

Modern neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the prosody of generated utterances often represents the average prosodic style of the database instead of having wide…

音频与语音处理 · 电气工程与系统科学 2020-09-16 Tuomo Raitio , Ramya Rasipuram , Dan Castellani

Multimodal sentiment analysis aims to identify the emotions expressed by individuals through visual, language, and acoustic cues. However, most existing research assume that all modalities are available during both training and testing,…

声音 · 计算机科学 2026-04-21 Weide Liu , Huijing Zhan
‹ 上一页 1 8 9 10 下一页 ›