中文
相关论文

相关论文: What Do Prosody and Text Convey? Characterizing Ho…

200 篇论文

Prominences and boundaries are the essential constituents of prosodic structure in speech. They provide for means to chunk the speech stream into linguistically relevant units by providing them with relative saliences and demarcating them…

计算与语言 · 计算机科学 2015-10-08 Antti Suni , Daniel Aalto , Martti Vainio

This dissertation proposes the study of multimodal learning in the context of musical signals. Throughout, we focus on the interaction between audio signals and text information. Among the many text sources related to music that can be used…

声音 · 计算机科学 2021-11-01 Gabriel Meseguer-Brocal

The differences in written text and conversational speech are substantial; previous parsers trained on treebanked text have given very poor results on spontaneous speech. For spoken language, the mismatch in style also extends to prosodic…

计算与语言 · 计算机科学 2020-10-12 Trang Tran , Jiahong Yuan , Yang Liu , Mari Ostendorf

Prosody plays a vital role in verbal communication. Acoustic cues of prosody have been examined extensively. However, prosodic characteristics are not only perceived auditorily, but also visually based on head and facial movements. The…

计算与语言 · 计算机科学 2022-09-14 Hartmut Meister , Isa Samira Winter , Moritz Waeachtler , Pascale Sandmann , Khaled Abdellatif

Humans communicate using systems of interconnected stimuli or concepts -- from language and music to literature and science -- yet it remains unclear how, if at all, the structure of these networks supports the communication of information.…

物理与社会 · 物理学 2020-03-27 Christopher W. Lynn , Lia Papadopoulos , Ari E. Kahn , Danielle S. Bassett

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on emotions, dialogues,…

声音 · 计算机科学 2025-06-27 Charuta Pethe , Bach Pham , Felix D Childress , Yunting Yin , Steven Skiena

Speech signals encode emotional, linguistic, and pathological information within a shared acoustic channel; however, disentanglement is typically assessed indirectly through downstream task performance. We introduce an information-theoretic…

声音 · 计算机科学 2026-02-25 Bipasha Kashyap , Björn W. Schuller , Pubudu N. Pathirana

Conversations emerge as the primary media for exchanging ideas and conceptions. From the listener's perspective, identifying various affective qualities, such as sarcasm, humour, and emotions, is paramount for comprehending the true…

计算与语言 · 计算机科学 2022-11-23 Shivani Kumar , Ishani Mondal , Md Shad Akhtar , Tanmoy Chakraborty

We investigate correlations in information carriers, e.g. texts and pieces of music, which are represented by strings of letters. For information carrying strings generated by one source (i.e. a novel or a piece of music) we find…

统计力学 · 物理学 2007-05-23 Werner Ebeling , Thorsten Poeschel , Karl-Friedrich Albrecht

Understanding and communicating data uncertainty is crucial for making informed decisions in sectors like finance and healthcare. Previous work has explored how to express uncertainty in various modes. For example, uncertainty can be…

人机交互 · 计算机科学 2024-04-15 Chase Stokes , Chelsea Sanker , Bridget Cogley , Vidya Setlur

The multimedia communications with texts and images are popular on social media. However, limited studies concern how images are structured with texts to form coherent meanings in human cognition. To fill in the gap, we present a novel…

多媒体 · 计算机科学 2023-02-28 Chunpu Xu , Hanzhuo Tan , Jing Li , Piji Li

Some recent models for Text-to-Speech synthesis aim to transfer the prosody of a reference utterance to the generated target synthetic speech. This is done by using a learned embedding of the reference utterance, which is used to condition…

计算与语言 · 计算机科学 2023-03-09 Atli Thor Sigurgeirsson , Simon King

In emotion recognition from speech, a key challenge lies in identifying speech signal segments that carry the most relevant acoustic variations for discerning specific emotions. Traditional approaches compute functionals for features such…

计算与语言 · 计算机科学 2025-06-04 Sofoklis Kakouros

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

计算与语言 · 计算机科学 2025-08-19 Shumin Que , Anton Ragni

Speech is a multiplexed signal displaying levels of complexity, organizational principles and perceptual units of analysis at distinct timescales. This critical acoustic signal for human communication is thus characterized at distinct…

神经元与认知 · 定量生物学 2024-07-10 Jérémy Giroud , Benjamin Morillon

Counseling is carried out as spoken conversation between a therapist and a client. The empathy level expressed by the therapist is considered an important index of the quality of counseling and often assessed by an observer or the client.…

音频与语音处理 · 电气工程与系统科学 2023-10-24 Dehua Tao , Tan Lee , Harold Chui , Sarah Luk

Recent advancements in neural end-to-end TTS models have shown high-quality, natural synthesized speech in a conventional sentence-based TTS. However, it is still challenging to reproduce similar high quality when a whole paragraph is…

声音 · 计算机科学 2022-09-15 Liumeng Xue , Frank K. Soong , Shaofei Zhang , Lei Xie

The prosody of a spoken word is determined by its surrounding context. In incremental text-to-speech synthesis, where the synthesizer produces an output before it has access to the complete input, the full context is often unknown which can…

计算与语言 · 计算机科学 2021-06-16 Brooke Stephenson , Thomas Hueber , Laurent Girin , Laurent Besacier

Language models have become nearly ubiquitous in natural language processing applications achieving state-of-the-art results in many tasks including prosody. As the model design does not define predetermined linguistic targets during…

计算与语言 · 计算机科学 2023-04-26 Sofoklis Kakouros , Johannah O'Mahony

Parsing spoken dialogue poses unique difficulties, including disfluencies and unmarked boundaries between sentence-like units. Previous work has shown that prosody can help with parsing disfluent speech (Tran et al. 2018), but has assumed…

计算与语言 · 计算机科学 2021-10-13 Elizabeth Nielsen , Mark Steedman , Sharon Goldwater