English
Related papers

Related papers: Gesture2Speech: How Far Can Hand Movements Shape E…

200 papers

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

Research in linguistics shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse. For example, speakers perform hand gestures to indicate topic shifts, helping listeners identify transitions in discourse. In…

Computation and Language · Computer Science 2025-03-06 Varsha Suresh , M. Hamza Mughal , Christian Theobalt , Vera Demberg

Gestures are inherent to human interaction and often complement speech in face-to-face communication, forming a multimodal communication system. An important task in gesture analysis is detecting a gesture's beginning and end. Research on…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Esam Ghaleb , Ilya Burenko , Marlou Rasenberg , Wim Pouw , Ivan Toni , Peter Uhrig , Anna Wilson , Judith Holler , Aslı Özyürek , Raquel Fernández

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational…

Sound · Computer Science 2023-05-04 Jinlong Xue , Yayue Deng , Fengping Wang , Ya Li , Yingming Gao , Jianhua Tao , Jianqing Sun , Jiaen Liang

The prosody of a spoken utterance, including features like stress, intonation and rhythm, can significantly affect the underlying semantics, and as a consequence can also affect its textual translation. Nevertheless, prosody is rarely…

Computation and Language · Computer Science 2024-11-01 Ioannis Tsiamas , Matthias Sperber , Andrew Finch , Sarthak Garg

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

Sound · Computer Science 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

Modern sequence to sequence neural TTS systems provide close to natural speech quality. Such systems usually comprise a network converting linguistic/phonetic features sequence to an acoustic features sequence, cascaded with a neural…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-26 Slava Shechtman , Alex Sorin

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

Machine Learning · Computer Science 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Evonne Ng , Shiry Ginosar , Trevor Darrell , Hanbyul Joo

When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Gwantae Kim , Seonghyeok Noh , Insung Ham , Hanseok Ko

We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual training objectives to…

Sound · Computer Science 2026-03-13 Suvendu Sekhar Mohanty

Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained…

Sound · Computer Science 2024-07-19 Weiqin Li , Peiji Yang , Yicheng Zhong , Yixuan Zhou , Zhisheng Wang , Zhiyong Wu , Xixin Wu , Helen Meng

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody given a single text input. While previous research has addressed this by predicting prosodic information from both text…

Computation and Language · Computer Science 2025-08-19 Shumin Que , Anton Ragni

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the starting point and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-04 Daria Diatlova , Vitaly Shutov

Generating gestures from human speech has gained tremendous progress in animating virtual avatars. While the existing methods enable synthesizing gestures cooperated by individual self-talking, they overlook the practicality of concurrent…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Xingqun Qi , Yatian Wang , Hengyuan Zhang , Jiahao Pan , Wei Xue , Shanghang Zhang , Wenhan Luo , Qifeng Liu , Yike Guo

Turn-taking is a fundamental aspect of human communication where speakers convey their intention to either hold, or yield, their turn through prosodic cues. Using the recently proposed Voice Activity Projection model, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Erik Ekstedt , Siyang Wang , Éva Székely , Joakim Gustafson , Gabriel Skantze

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

In recent years because of the advances in computer vision research, free hand gestures have been explored as means of human-computer interaction (HCI). Together with improved speech processing technology it is an important step toward…

Computer Vision and Pattern Recognition · Computer Science 2007-05-23 S. Kettebekov , R. Sharma

Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utterance. Compared to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Muhammad Hamza Mughal , Rishabh Dabral , Ikhsanul Habibie , Lucia Donatelli , Marc Habermann , Christian Theobalt

Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-05 Cedric Chan , Jianjing Kuang
‹ Prev 1 2 3 10 Next ›