中文
相关论文

相关论文: What Do Prosody and Text Convey? Characterizing Ho…

200 篇论文

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

声音 · 计算机科学 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model based on a conditional…

声音 · 计算机科学 2024-12-05 Jiaxuan Liu , Zhaoci Liu , Yajun Hu , Yingying Gao , Shilei Zhang , Zhenhua Ling

Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet evaluation metrics remain blind to such phenomena. We meta-evaluate both text- and speech-based…

计算与语言 · 计算机科学 2026-05-28 Maike Züfle , Danni Liu , Vilém Zouhar , Jan Niehues

Conversations contain a wide spectrum of multimodal information that gives us hints about the emotions and moods of the speaker. In this paper, we developed a system that supports humans to analyze conversations. Our main contribution is…

人机交互 · 计算机科学 2020-01-29 Joshua Y. Kim , Greyson Y. Kim , Kalina Yacef

We study the information transmission through a quantum channel, defined over a continuous alphabet and losing its energy en route, in presence of correlated noise among different channel uses. We then show that entangled inputs improve the…

量子物理 · 物理学 2009-11-11 Giovanna Ruggeri , Giulio Soliani , Vittorio Giovannetti , Stefano Mancini

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…

Expressive human speech generally abounds with rich and flexible speech prosody variations. The speech prosody predictors in existing expressive speech synthesis methods mostly produce deterministic predictions, which are learned by…

声音 · 计算机科学 2023-10-10 Xiang Li , Songxiang Liu , Max W. Y. Lam , Zhiyong Wu , Chao Weng , Helen Meng

While sentiment and emotion analysis have been studied extensively, the relationship between sarcasm and emotion has largely remained unexplored. A sarcastic expression may have a variety of underlying emotions. For example, "I love being…

计算与语言 · 计算机科学 2022-06-07 Anupama Ray , Shubham Mishra , Apoorva Nunna , Pushpak Bhattacharyya

Human communication involves more than explicit semantics, with implicit signals and contextual cues playing a critical role in shaping meaning. However, modern speech technologies, such as Automatic Speech Recognition (ASR) and…

Whether animal or speech communication, environmental sounds, or music -- all sounds carry some information. Sound sources are embedded in acoustic environments that contain any number of additional sources that emit sounds that reach the…

神经元与认知 · 定量生物学 2019-02-21 Adam Weisser

Conversation has been a primary means for the exchange of information since ancient times. Understanding patterns of information flow in conversations is a critical step in assessing and improving communication quality. In this paper, we…

计算与语言 · 计算机科学 2021-09-15 Laurence A. Clarfeld , Robert Gramling , Donna M. Rizzo , Margaret J. Eppstein

We present a document compression system that uses a hierarchical noisy-channel model of text production. Our compression system first automatically derives the syntactic structure of each sentence and the overall discourse structure of the…

计算与语言 · 计算机科学 2009-07-07 Hal Daumé , Daniel Marcu

We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use either fixed-dimensional or variable-length prosody embedding via…

音频与语音处理 · 电气工程与系统科学 2019-07-05 Viacheslav Klimkov , Srikanth Ronanki , Jonas Rohnke , Thomas Drugman

Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing…

机器学习 · 计算机科学 2024-08-23 Luyao Cheng , Hui Wang , Siqi Zheng , Yafeng Chen , Rongjie Huang , Qinglin Zhang , Qian Chen , Xihao Li

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable…

音频与语音处理 · 电气工程与系统科学 2022-02-17 Yi Ren , Ming Lei , Zhiying Huang , Shiliang Zhang , Qian Chen , Zhijie Yan , Zhou Zhao

Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods relying on speech prompts (reference speech) for voice…

音频与语音处理 · 电气工程与系统科学 2023-10-13 Yichong Leng , Zhifang Guo , Kai Shen , Xu Tan , Zeqian Ju , Yanqing Liu , Yufei Liu , Dongchao Yang , Leying Zhang , Kaitao Song , Lei He , Xiang-Yang Li , Sheng Zhao , Tao Qin , Jiang Bian

Human language has a distinct systematic structure, where utterances break into individually meaningful words which are combined to form phrases. We show that natural-language-like systematicity arises in codes that are constrained by a…

计算与语言 · 计算机科学 2025-11-19 Richard Futrell , Michael Hahn

Speakers usually adjust their way of talking in noisy environments involuntarily for effective communication. This adaptation is known as the Lombard effect. Although speech accompanying the Lombard effect can improve the intelligibility of…

音频与语音处理 · 电气工程与系统科学 2019-04-10 Yi Zhao , Atsushi Ando , Shinji Takaki , Junichi Yamagishi , Satoshi Kobashikawa

Sarcasm is a rhetorical device that is used to convey the opposite of the literal meaning of an utterance. Sarcasm is widely used on social media and other forms of computer-mediated communication motivating the use of computational models…

计算与语言 · 计算机科学 2024-10-25 Shafkat Farabi , Tharindu Ranasinghe , Diptesh Kanojia , Yu Kong , Marcos Zampieri

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

音频与语音处理 · 电气工程与系统科学 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi